論文の概要: Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement
- arxiv url: http://arxiv.org/abs/2609.01148v1
- Date: Tue, 01 Sep 2026 12:24:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.645341
- Title: Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement
- Title(参考訳): Dotting the Eye:視覚強調のためのインテント駆動型画像修正エージェント
- Authors: Chujie Qin, Zilong Zhang, Zewei Chang, Chunle Guo, Ruixing Wang, Tao Hu, Ming-Ming Cheng, Chongyi Li,
- Abstract要約: MLLM駆動のリタッチエグゼキュータであるEyeControlを提案する。
わずか数クリックまたは粗いストロークで、EyeControlは意図した領域に視覚的注意を向け、画像の事実上の「目を引く」。
また,グローバルな調整と局所的な調整のコーディネーションを改善するための操作一貫性制約を導入し,より自然でコヒーレントなリタッチを実現する。
- 参考スコア(独自算出の注目度): 85.49085209742798
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Image retouching is commonly formulated as enhancing overall visual quality through color adjustment, but in practice, it also serves to emphasize visual focus by guiding viewers' attention toward a specific subject or region. Achieving such focus-oriented retouching is inherently challenging, as it requires well-coordinated global and local adjustments to manipulate perceptual saliency while maintaining visual naturalness. This intricate process typically demands substantial professional expertise. In this study, we propose EyeControl, a MLLM-driven agent with a diffusion-based retouching executor that enables visual focus enhancement under weak user intent. With only a few clicks or coarse strokes, EyeControl directs visual attention to the intended region, effectively "dotting the eye" of the image. The core idea is to explicitly link the weak user intention with the target editing region and the corresponding tonal adjustment operations during retouching. To achieve this, the system first interprets the intent and image content to infer the visual focus and generate structured intent guidance for the retouching executor. Second, the retouching executor is encouraged to respond more strongly to the target region, explicitly aligning its attention map with a designed pseudo-intent map. We also introduce an operation-consistency constraint to improve coordination between global and local adjustments, achieving more natural and coherent retouching. Additionally, we contribute ControlArt-Bench, a high-quality evaluation dataset for visual focus enhancement. Extensive evaluations demonstrate that EyeControl yields perceptually appealing results with stronger intent alignment. Code can be found in https://github.com/DragonisCV/EyeControl.
- Abstract(参考訳): 画像のリタッチは、色調整による全体的な視覚的品質の向上として一般的に定式化されているが、実際には、特定の主題や領域に対して視聴者の注意を向けることによる視覚的焦点の強調にも役立っている。
このようなフォーカス指向のリタッチを実現することは本質的に困難であり、視覚的自然性を維持しながら知覚の正当性を操作するために、協調したグローバルな調整と局所的な調整が必要である。
この複雑なプロセスは通常、相当な専門知識を必要とします。
本研究では,ユーザ意図の弱い視覚的焦点強調を可能にする拡散型リタッチエグゼキュータを用いたMLLMエージェントであるEyeControlを提案する。
わずか数クリックまたは粗いストロークで、EyeControlは意図した領域に視覚的注意を向け、画像の事実上の「目を引く」。
中心となる考え方は、弱いユーザ意図とターゲット編集領域と、リタッチ中の対応する音調調整操作とを明示的に関連付けることである。
これを実現するために、まずインテントと画像の内容を理解して視覚的焦点を推測し、リタッチ実行者のための構造化されたインテントガイダンスを生成する。
第2に、リタッチ実行者は、ターゲット領域に対してより強く応答することを奨励し、その注意マップを設計された疑似意図マップに明示的に整合させる。
また,グローバルな調整と局所的な調整のコーディネーションを改善するための操作一貫性制約を導入し,より自然でコヒーレントなリタッチを実現する。
また、視覚焦点強調のための高品質な評価データセットであるControlArt-Benchにコントリビュートする。
広範囲な評価は、EyeControlがより強い意図的アライメントで知覚的に魅力的な結果をもたらすことを示している。
コードはhttps://github.com/DragonisCV/EyeControlにある。
関連論文リスト
- Tinted Frames: Question Framing Blinds Vision-Language Models [29.78944164519993]
VLM(Vision-Language Models)は、視覚的推論を必要とするタスクでも視覚的な入力をあまり利用していないことが示されている。
我々は、フレーミングが画像上の注意の量と分布の両方を変えるかを定量化する。
本稿では,学習可能なトークンを用いた軽量なプロンプトチューニング手法を提案する。
論文 参考訳(メタデータ) (2026-03-19T17:53:09Z) - A Saccade-inspired Approach to Image Classification using Vision Transformer Attention Maps [0.9332987715848716]
人間の視覚システムからインスピレーションを得て、よりスマートな画像処理モデルを作成します。
自己教師型視覚変換器であるDINOを用いて,視覚空間の重要領域に情報処理を集中させるササードインスピレーション方式を提案する。
この選択的処理戦略は、フルイメージの分類性能の大部分を保ち、場合によっては性能も向上する。
論文 参考訳(メタデータ) (2026-03-10T12:54:55Z) - PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching [54.3683137773426]
本稿ではPerTouchと呼ばれる拡散型画像修正フレームワークを提案する。
本手法は,グローバルな美学を維持しつつ,セマンティックレベルのイメージリタッチをサポートする。
我々は,強力なユーザ命令と弱いユーザ命令の両方を扱えるVLMエージェントを開発した。
論文 参考訳(メタデータ) (2025-11-17T05:39:15Z) - RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models [76.79706360982162]
トレーニング不要なホワイトボックス画像リタッチシステムであるRetouchLLMを提案する。
高解像度の画像に直接、解釈可能でコードベースのリタッチを実行する。
我々のフレームワークは、人間がマルチステップのリタッチを行う方法と同じような方法で、徐々に画像を強化する。
論文 参考訳(メタデータ) (2025-10-09T10:40:49Z) - Top-Down Visual Attention from Analysis by Synthesis [87.47527557366593]
我々は、古典的分析・合成(AbS)の視覚的視点からトップダウンの注意を考察する。
本稿では,AbSを変動的に近似したトップダウン変調ViTモデルであるAbSViT(Analytic-by-Synthesis Vision Transformer)を提案する。
論文 参考訳(メタデータ) (2023-03-23T05:17:05Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。