論文の概要: IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
- arxiv url: http://arxiv.org/abs/2604.05157v1
- Date: Mon, 06 Apr 2026 20:39:30 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-08 17:42:09.48045
- Title: IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
- Title(参考訳): IntentScore:コンピュータ・ユース・エージェントに対するインテント・コンディションド・アクション・アセスメント
- Authors: Rongqian Chen, Yu Li, Zeyu Fang, Sizhe Tang, Weidong Cao, Tian Lan,
- Abstract要約: IntentScoreは、398KオフラインGUIインタラクションステップから候補動作のスコアを学習するプラン対応報酬モデルである。
Int IntentScore 97.5%は、ホールドアウト評価においてペアワイズ判別精度を達成する。
Int IntentScoreはタスク成功率を6.9ポイント改善し、不均一なオフライン軌道から学んだ報酬推定が未確認エージェントやタスク分布に一般化されることを示した。
- 参考スコア(独自算出の注目度): 10.905829987425752
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evaluating action quality, leading to irreversible errors that cascade through subsequent steps. We propose IntentScore, a plan-aware reward model that learns to score candidate actions from 398K offline GUI interaction steps spanning three operating systems. IntentScore trains with two complementary objectives: contrastive alignment for state-action relevance and margin ranking for action correctness. Architecturally, it embeds each candidate's planning intent in the action encoder, enabling discrimination between candidates with similar actions but different rationales. IntentScore achieves 97.5% pairwise discrimination accuracy on held-out evaluation. Deployed as a re-ranker for Agent S3 on OSWorld, an environment entirely unseen during training, IntentScore improves task success rate by 6.9 points, demonstrating that reward estimation learned from heterogeneous offline trajectories generalizes to unseen agents and task distributions.
- Abstract(参考訳): Computer-Use Agents (CUA) は、大規模な言語モデルを利用してデスクトップ環境でGUI操作を実行するが、アクション品質を評価せずにアクションを生成し、その後のステップでカスケードする不可逆的なエラーを引き起こす。
IntentScoreは、3つのオペレーティングシステムにまたがる398KオフラインGUIインタラクションステップから、候補動作のスコアを学習するプラン対応報酬モデルである。
IntentScoreは2つの相補的な目的を持つ列車である。
アーキテクチャ的には、各候補の計画意図をアクションエンコーダに組み込んで、類似のアクションを持つ候補同士の識別を可能にする。
IntentScoreはホールドアウト評価において97.5%のペアワイズ判別精度を達成する。
IntentScoreはOSWorld上のエージェントS3のリランカとしてデプロイされ、トレーニング中に完全に見えない環境であり、タスク成功率を6.9ポイント改善し、不均一なオフライン軌跡から学んだ報酬推定が未確認エージェントやタスク配布に一般化することを示した。
関連論文リスト
- IntentCUA: Learning Intent-level Representations for Skill Abstraction and Multi-Agent Planning in Computer-Use Agents [4.655926959889001]
IntentCUAは,計画メモリによる長期実行の安定化を目的としたコンピュータ用フレームワークである。
Int Intentプロトタイプはサブグループ対応のスキルを取得し、部分的な計画にそれらを注入することで、冗長な再計画が削減される。
Int IntentCUAは、ステップ効率比0.91で74.83%のタスク成功率を達成した。
論文 参考訳(メタデータ) (2026-02-19T03:42:15Z) - When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents [50.5814495434565]
この研究は、コンピュータ利用エージェント(CUA)における不整合検出を定義し、研究する最初の試みである。
実世界のCUAデプロイメントにおける3つの一般的なカテゴリを特定し、人間の注釈付きアクションレベルのアライメントラベルを用いたリアルな軌跡のベンチマークであるMisActBenchを構築した。
本稿では,実行前に不整合を検知し,構造化されたフィードバックによって繰り返し修正する,実用的で普遍的なガードレールであるDeActionを提案する。
論文 参考訳(メタデータ) (2026-02-09T18:41:15Z) - GTA1: GUI Test-time Scaling Agent [97.58177633084915]
グラフィカルユーザインタフェース(GUI)は、ユーザ命令をアクションプロポーザルに順次分解することで、プラットフォーム(例えばLinux)間で自律的にタスクを完了させる。
本稿では,前述の textbfGUI textbfTest-time Scaling textbfAgent,すなわち GTA1 の課題について検討する。
論文 参考訳(メタデータ) (2025-07-08T08:52:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。