論文の概要: AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
- arxiv url: http://arxiv.org/abs/2609.35530v1
- Date: Mon, 28 Sep 2026 16:13:37 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 21:04:19.734028
- Title: AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
- Title(参考訳): AutoRef: エージェントマルチ参照画像生成のためのハーネス最適化
- Abstract要約: 両モデルの凍結を維持しながら,ハーネスを自動的に最適化するAutoRefを提案する。
この手法を用いて,オープンウェイトFLUX.2[klein]4Bを5.72から7.37に改善したAutoRef-Harnessを,ホールドアウト4参照タスクで発見する。
再最適化なしでは、ジェネレータ、参照数、ベンチマーク、評価器、推論モデルが検索で使用されるものと異なる場合、同じハーネスが結果を改善する。
- 参考スコア(独自算出の注目度): 40.132233802644244
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are interpreted, how generation is performed, how outputs are diagnosed, and how the final image is selected. In multi-reference generation, however, references play different roles and outputs must satisfy many criteria at once, such as fidelity to each reference and the naturalness of the whole image, so many parts of the harness could be improved, from how references are processed to how outputs are diagnosed. This makes it hard to predict which changes will improve performance and by how much, and good harnesses difficult to design by hand; indeed, human-written harnesses vary widely in performance. We therefore propose AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. AutoRef separates the tasks whose feedback informs proposals from the tasks used to select candidates, and continues the search from a beam of the top-ranked harnesses on the selection tasks. Using this procedure, we discover AutoRef-Harness, which improves the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models including Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improves results when the generator, number of references, benchmark, evaluator, or reasoning model differs from those used in the search.
- Abstract(参考訳): 最近の画像生成モデルは、複数の参照画像を入力として取り込み、それらを新しい画像に組み合わせることができる。
しかし、マルチ参照画像生成は依然として困難であり、モデルでは参照から被写体を省略または重複させたり、複数の被写体が不自然にペーストされた画像を生成したりすることができる。
近年、画像生成モデル、推論モデル、ハーネスを組み合わせた画像生成エージェントが提案されている。これは、参照画像の解釈方法、生成方法、出力の診断方法、最終的な画像の選択方法などを特定する実行可能なプログラムである。
しかし、マルチ参照生成では、参照は異なる役割を担い、出力は各参照への忠実さや画像全体の自然さといった多くの基準を一度に満たさなければならないため、参照がどのように処理されるかから出力の診断方法まで、ハーネスの多くの部分を改善することができる。
これにより、どの変更がパフォーマンスを向上し、どれだけの量で良いハーネスを手作業で設計することが難しいかを予測しにくくなります。
そこで我々は、両方のモデルを凍結させながらハーネスを自動的に最適化するAutoRefを提案する。
AutoRefは、フィードバックが候補の選択に使用されるタスクから提案を通知するタスクを分離し、選択タスクのトップランクのハーネスのビームからの検索を継続する。
本手法を用いて,Nano Banana Pro や GPT-Image-1.5 などのプロプライエタリモデルに適合または超過したマルチバナベンチマークの4参照タスクにおいて,オープンウェイトFLUX.2 [klein] 4B を 5.72 から 7.37 に改善した AutoRef-Harness が発見された。
再最適化なしでは、ジェネレータ、参照数、ベンチマーク、評価器、推論モデルが検索で使用されるものと異なる場合、同じハーネスが結果を改善する。
関連論文リスト
- RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing [9.858046310638803]
本稿では,評価次元,重み,評価基準を有する入力固有ルーリックを生成し,そのルーリックを候補画像に応用するペアワイズ生成報酬モデリングフレームワークを提案する。
我々は、2段階のトレーニングパイプラインを用いて、テキスト・ツー・イメージ生成と画像編集のための専用RMモデルを訓練する。
論文 参考訳(メタデータ) (2026-08-27T10:56:47Z) - GMAIL: Generative Modality Alignment for generated Image Learning [51.071351994330605]
本稿では,生成画像の識別のための新しいフレームワークGMAILを提案する。
我々のフレームワークは様々な視覚言語モデルに容易に組み込むことができ、広範囲にわたる実験を通してその有効性を示す。
論文 参考訳(メタデータ) (2026-02-17T05:40:25Z) - ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation [25.39019070750831]
拡散モデルは珍しい概念や目に見えない概念を生み出すのに苦労する。
本稿では,あるテキストプロンプトに基づいて関連画像を動的に検索するImageRAGを提案する。
私たちのアプローチは高度に適応可能で、異なるモデルタイプにまたがって適用できます。
論文 参考訳(メタデータ) (2025-02-13T15:36:12Z) - Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers [57.62831463679979]
本稿では,近年の離散最適化手法の突発的逆転問題に対する直接比較について述べる。
逆プロンプトと基底真理画像とのCLIP類似性に着目し, 逆プロンプトが生成する画像と基底真理画像との類似性について検討した。
論文 参考訳(メタデータ) (2024-08-12T21:35:59Z) - Re-Imagen: Retrieval-Augmented Text-to-Image Generator [58.60472701831404]
検索用テキスト・ツー・イメージ・ジェネレータ(再画像)
検索用テキスト・ツー・イメージ・ジェネレータ(再画像)
論文 参考訳(メタデータ) (2022-09-29T00:57:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。