FuguReport

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

Authors Zekai Zhang, Jiahao Li, Jie Zhang, Kaiyuan Gao, Kun Yan, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiaoyue Chen, Xiao Xu, Yan Shu, Yanran Zhang, Yixian Xu, Yuxiang Chen, Zhendong Wang, Zihao Liu, Zikai Zhou, Huishuai Zhang, Dongyan Zhao, Chenfei Wu
Categories Method / Context-aware Agent / Integrating planning, reasoning, and feedback, Application / Image Generation / Real-world image generation with context, Evaluation / Model Evaluation / Performance benchmarking on IA-Bench and others
License CC BY 4.0

Abstract Overview

This paper frames a central limitation of real-world text-to-image generation as a "context gap" between the partial information users provide and the fuller generation context a model actually needs. To address this, it proposes Qwen-Image-Agent, a unified agentic framework that incrementally constructs generation context through Context-Aware Planning and Context Grounding. The planning module operates at information, content, and generation levels, while grounding gathers missing context from reasoning, search, memory, and feedback. The paper also introduces IA-Bench, a benchmark designed to evaluate agentic image generation across planning, reasoning, search, and memory in real-world tasks.

Novelty

The work’s main novelty is to treat real-world image generation explicitly as a context-construction problem rather than direct prompt rendering. It combines planning, reasoning, search, memory, and feedback in a single context-centric framework and pairs this with IA-Bench, a benchmark that evaluates these agentic capabilities jointly rather than in isolation.

Results

On IA-Bench, Qwen-Image-Agent achieves the highest reported IA-score of 45.4, exceeding strong closed-source baselines and improving substantially over the direct Qwen-Image-2.0 baseline of 17.4. It also reports state-of-the-art results on WISE-Verified (overall 0.9020) and MindBench (overall 0.42). Ablation studies show that removing reason, search, memory, or feedback degrades the corresponding capability, supporting the claimed contribution of grounded context construction.

Key Points

  1. Qwen-Image-Agent builds generation context progressively from partial user context using context-aware planning and grounding via reason, search, memory, and feedback.
  2. The paper introduces IA-Bench, which covers four agentic capabilities across 17 subtasks, 730 instances, and 1801 checklist items for structured evaluation.
  3. Empirically, the framework establishes state-of-the-art performance on IA-Bench, WISE-Verified, and MindBench, with particularly large gains over direct generation methodologies.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.