論文の概要: Agent Skills Should Go Beyond Text: The Case for Visual Skills
- arxiv url: http://arxiv.org/abs/2606.01414v1
- Date: Sun, 31 May 2026 19:22:43 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-02 21:34:29.691207
- Title: Agent Skills Should Go Beyond Text: The Case for Visual Skills
- Title(参考訳): エージェントスキルはテキストを超えるべき:ビジュアルスキルのケース
- Authors: Binxiao Xu, Ruichuan An, Bocheng Zou, Hang Hua,
- Abstract要約: 再利用可能なスキルは、エージェント能力を拡張するための重要なメカニズムである。
既存のスキル学習手法の多くは、再利用可能な体験をテキストのみの資産として保存する。
このテキストのみのパラダイムは、視覚中心のタスクに根本的なボトルネックをもたらすと我々は主張する。
- 参考スコア(独自算出の注目度): 9.601673227260926
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Reusable skills are a key mechanism for extending agent capabilities, allowing agents to accumulate experience and solve increasingly complex tasks. Yet most existing skill-learning methods store reusable experience as text-only assets, such as instructions, reasoning traces, or summarized trajectories. We argue that this text-only paradigm creates a fundamental bottleneck for visual-centric tasks, where reusable knowledge often depends on spatial layout, visual grounding, fine-grained appearance, and localized state changes. To address this limitation, we propose \textbf{\NAME}, a multimodal skill paradigm that combines declarative textual logic with explicit visual support. We distinguish three reusable forms: static priors for stable spatial conventions, dynamic priors for in-situ visual working memory, and interleaved visual skills that bind ordered text steps to the source frames, screenshots, or page regions that justify them. Rather than only describing what to do, visual skills also encode where to look, how to inspect, and how to verify visual outcomes. To scale visual-skill construction, we introduce \textbf{\SYSTEM}, an automatic system that converts agent experience into reusable multimodal skills by preserving textual reasoning, spatial references, visual boundaries, and interaction patterns from task trajectories. Experiments on GUI and other visual-centric tasks show that visual skills consistently outperform text-only skills, particularly when success requires spatial correspondence, visual evidence, and state-aware interaction. These results support our central position: reusable agent skills should go beyond text and become multimodal assets for future multimodal agents.
- Abstract(参考訳): 再利用可能なスキルは、エージェント能力を拡張するための重要なメカニズムであり、エージェントが経験を蓄積し、ますます複雑なタスクを解決することができる。
しかし、既存のスキル学習手法の多くは、再利用可能な経験を、指示、推論トレース、要約された軌跡など、テキストのみの資産として保存している。
このテキストのみのパラダイムは、再利用可能な知識は、しばしば空間的レイアウト、視覚的接地、きめ細かい外観、局所的な状態変化に依存する。
この制限に対処するために,宣言型テキスト論理と明示的な視覚的サポートを組み合わせたマルチモーダルスキルパラダイムである‘textbf{\NAME} を提案する。
安定した空間的慣行の静的先行、その場での視覚的作業メモリの動的先行、およびそれらを正当化するソースフレーム、スクリーンショット、ページ領域に順序付きテキストステップをバインドするインターリーブ視覚スキルの3つの再利用可能な形式を区別する。
視覚的なスキルは、何をすべきかを説明するだけでなく、見るべき場所、検査方法、そして視覚的な結果の検証方法をコード化します。
本稿では,タスク軌跡からテキスト推論,空間参照,視覚的境界,インタラクションパターンを保存することで,エージェント体験を再利用可能なマルチモーダルスキルに変換するシステムである‘textbf{\SYSTEM}を紹介する。
GUIやその他の視覚中心のタスクの実験は、特に成功が空間的対応、視覚的証拠、状態認識相互作用を必要とする場合、視覚スキルがテキストのみのスキルより一貫して優れていることを示している。
再利用可能なエージェントスキルは、テキストを超えて、将来のマルチモーダルエージェントのためのマルチモーダルアセットになるべきです。
関連論文リスト
- Visual Text Processing: A Comprehensive Review and Unified Evaluation [99.57846940547171]
視覚テキスト処理における最近の進歩を包括的・多視点的に分析する。
本研究の目的は,視覚テキスト処理のダイナミックな分野における今後の探索と革新を促進する基礎資源として,本研究を確立することである。
論文 参考訳(メタデータ) (2025-04-30T14:19:29Z) - Enhancing Visual Representation for Text-based Person Searching [9.601697802095119]
VFE-TPSは、ビジュアルフィーチャ強化テキストベースのPerson Searchモデルである。
基本的なマルチモーダル機能を学ぶために、トレーニング済みのバックボーンCLIPを導入する。
Text Guided Masked Image Modelingタスクを構築し、局所的な視覚的詳細を学習するモデルの能力を強化する。
論文 参考訳(メタデータ) (2024-12-30T01:38:14Z) - Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning [2.401993998791928]
本稿では、モダリティを接続するための軽量な視覚言語マッピングネットワークを訓練するフレームワークを提案する。
視覚的関連性やストーリー情報性も向上するマルチモーダルなコントラスト目標を提案する。
論文 参考訳(メタデータ) (2024-08-12T16:15:32Z) - Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model [25.47573567479831]
本稿では,視覚とテキストの両方のプロンプト技術を利用した新しい推論に基づく視覚的ICL手法を提案する。
提案手法はアウト・オブ・ボックスであり,微調整や最適化は不要である。
論文 参考訳(メタデータ) (2024-05-16T17:59:21Z) - Improving In-Context Learning in Diffusion Models with Visual
Context-Modulated Prompts [83.03471704115786]
本研究では,改良型プロンプト拡散(iPromptDiff)を紹介する。
iPromptDiffは、視覚コンテキストを埋め込みベクトルに変換するエンドツーエンドのトレーニングされた視覚エンコーダを統合する。
拡散に基づく視覚基盤モデルにおいて,この視覚的文脈変調テキストガイダンスと標準制御ネット構造を組み込んだ場合,多種多様な学習課題における多目的性と堅牢性を示すことを示す。
論文 参考訳(メタデータ) (2023-12-03T14:15:52Z) - Visually-augmented pretrained language models for NLP tasks without
images [77.74849855049523]
既存のソリューションはしばしば視覚的知識増強のために明示的なイメージに依存している。
我々は、新しいtextbfVisually-textbfAugmented fine-tuningアプローチを提案する。
我々のアプローチは、BERT、RoBERTa、BART、T5を異なるスケールで継続的に改善することができる。
論文 参考訳(メタデータ) (2022-12-15T16:13:25Z) - Leveraging Visual Knowledge in Language Tasks: An Empirical Study on
Intermediate Pre-training for Cross-modal Knowledge Transfer [61.34424171458634]
視覚的知識を言語モデルに組み込むことがギャップを埋めるかどうかを検討する。
実験の結果,視覚的知識伝達は低リソース環境と完全教師付き環境の両方で性能を向上できることがわかった。
論文 参考訳(メタデータ) (2022-03-14T22:02:40Z) - External Knowledge Augmented Text Visual Question Answering [0.6445605125467573]
本稿では,視覚言語理解タスクのための標準マルチモーダルトランスフォーマー上で知識を抽出,フィルタリング,エンコードするフレームワークを提案する。
2つの公開データセット上で、最先端のデータセットに匹敵する結果を生成する。
論文 参考訳(メタデータ) (2021-08-22T13:21:58Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。