FuguReport

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

Authors Chonghuinan Wang, Zhikai Chen, Chunwei Wang, Yecong Wan, Junwei Yang, Zhixin Wang, Wei Zhang, Jiaqi Xu, Renjing Pei, Xiaohe Wu, Fan Li, Wangmeng Zuo
Affiliations Harbin Institute of Technology / Nankai University / Huawei
Categories Method / Multimodal Models / Unified text-image generation paradigm, Method / Training Techniques / Progressive adaptive training objectives, Evaluation / Multimodal Evaluation / Objective evaluation for interleaved sequences
License CC BY 4.0

Abstract Overview

ILLUME-X is a unified multimodal model designed for free-form interleaved text-image generation, where text and images can appear in arbitrary sequences. The paper attributes its approach to three main components: a curated 100K-sample interleaved training pipeline, a progressive training strategy with self-adaptive objectives for variable-length multimodal token sequences, and a dedicated evaluation framework called ILScore. Architecturally, the model uses a shared-transformer design with modality-specific processing and an interleaved attention mechanism to jointly support text generation, image synthesis, and modality transitions. The experiments evaluate both interleaved generation and standard text-to-image generation, with comparisons against prior unified models and selected commercial or agent-based systems.

Novelty

The paper’s main distinction is its attempt to treat free-form N-to-M interleaved text-image generation as a unified modeling problem supported by coordinated advances in data curation, training, and evaluation. It also introduces ILScore, a metric specifically structured to assess interleaved sequences across image-text alignment, single-image quality, sequence consistency, and text quality.

Results

On ISG-Bench, ILLUME-X reports an average score of 6.26, which the authors describe as state-of-the-art among unified models and equal to the agent-based ISG-AGENT system. Under the proposed ILScore, it achieves the highest overall aggregate score of 5.34, narrowly above Emu 3.5 at 5.33. The model also reports strong text-to-image results, including 0.85 overall on GenEval and 86.38 overall on DPG-Bench, while using significantly less inference time per image than Emu 3.5.

Key Points

  1. ILLUME-X combines a 100K-sample interleaved data curation pipeline with progressive multimodal training to improve free-form text-image sequence generation.
  2. The method introduces interleaved attention and classifier-free guidance adaptations to handle arbitrary modality transitions within a unified transformer framework.
  3. The evaluation contribution, ILScore, is designed to measure interleaved generation more comprehensively than prior benchmark protocols, and experiments show competitive or leading performance on both interleaved and text-to-image benchmarks.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.