Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
Abstract
005 We present Lunima-OmniLV (abbreviated as OmniLV), 005006 a universal multimodal multi-task framework for low-level vision that ad- 006007 dresses over 100 sub-tasks across four major categories, including image 007008 restoration, image enhancement, weak-semantic dense prediction, and 008009 stylization. OmniLV leverages both textual and visual prompts to offer 009010 flexible, user-friendly interactions. Built on Diffusion Transformer (DiT)- 010011 based generative priors, our framework supports arbitrary resolutions 011012 — achieving optimal performance at 1K resolution — while preserving 012013 fine-grained details and high fidelity. Through extensive experiments, 013014 we demonstrate that separately encoding text and visual instructions, 014015 combined with co-training using shallow feature control, is essential to 015016 mitigate task ambiguity and enhance multi-task generalization. Our find- 016017 ings also reveal that integrating high-level generative tasks into low-level 017018 vision models can compromise detail-sensitive restoration. These insights 018019 pave the way for more robust and generalizable low-level vision systems. 019