Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
Abstract
The advancement of generative AI models capable of pro-ducing text and image marks a critical step forward in the realm of* Equal Contribution, B Corresponding Author, † Project Leadermultimodal intelligence, particularly for tasks involving the interleavingof both modalities. To advance this intelligence to the next stage, it iscrucial for models to autonomously generate free-form interleaved text-image sequences. In this paper, we introduce ILLUME-X, an advancedunified multimodal paradigm that enables high-quality, free-form inter-leaved text-image generation by improving multimodal data efficiencyand stabilizing the multimodal training process. ILLUME-X comprisesthree key components: (i) an expanded training data pipeline optimizedfor interleaved text-image generation, (ii) a progressive training strategywith self-adaptive objectives for free-length multimodal token sequences,and (iii) an objective and comprehensive evaluation method ILScore forinterleaved text-image sequences. Notably, our ILLUME-X outperformsprevious unified models across multiple interleaved text-image generationtasks like style transfer, image decomposition and storytelling.