Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
Abstract
Unified Multimodal Models (UMMs) excel in text-to-image generation and editing but often entangle multimodal generative reasoning with high-fidelity visual synthesis. We introduce Query-Kontext, a novel approach that bridges a Vision-Language Model (VLM) and a diffusion model via multimodal “kontext” tokens. This design cleanly decouples the complex reasoning delegated to the VLM, including instruction understanding, grounding, and identity preservation, from the high-quality visual rendering executed by the diffusion model. We propose a three-stage progressive training strategy: (1) connecting the VLM to a lightweight diffusion head to activate generative reasoning; (2) scaling to a large, pre-trained diffusion model to enhance visual realism; and (3) incorporating a low-level image encoder for fine-grained instruction tuning. Supported by a comprehensive multi-task dataset, extensive experiments demonstrate that Query-Kontext matches or outperforms state-of-theart task-specific and unified methods across diverse reference-to-image scenarios.