論文の概要: Teach it to stop, not just to click
- arxiv url: http://arxiv.org/abs/2607.17136v1
- Date: Sun, 19 Jul 2026 08:46:06 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.378718
- Title: Teach it to stop, not just to click
- Title(参考訳): 単にクリックするだけでなく、止まるように教える
- Abstract要約: 修復された政策の成功率は上流の足場に支配されていることを示す。
通常のk-seedレポートのためのライブラリ(cua_reliability)をリリースする。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Agentic computer-use RL is reported in single runs, and those numbers mislead. Using verifier-guided repair of a 35B computer-use agent (CUA) across five oracle-graded environments, we show a repaired policy's success rate is dominated by upstream variance: a variance-components decomposition across three cells (crossed data-draw $\times$ seed grid, bootstrap CIs) finds evaluation variance negligible ($σ_{\mathrm{eval}} \approx 0$) and the training-seed effect small everywhere ($\leq 10\%$); instead it splits between the data draw and run-to-run nondeterminism, the data draw's share rising to dominant ($48\%$) on the hardest cell. There the run-to-run distribution is bimodal (Hartigan dip $p=0.07$, $k=10$), so a single run has roughly a 30% chance of the failure mode and mean$\pm$std is the wrong summary. On that footing, two findings hold. First, repairability is two-tier in how constrained the corrective action is: a single fixed token installs reliably (done-detection $0.97\pm0.06$), while open-ended corrections are only partial -- spatial-coordinate clicks (grounding $0.53\pm0.35$) and a generative field-fill ($0.14\pm0.04$). Second, the frame-level repair transfers to task success only when the corrective action is the task's sole remaining blocker (LinkedIn 8/20 vs. base 0/15, Fisher $p=0.006$). We caught two of our own over-claims -- a sample-efficiency curve and a 'grounding cannot be bought' boundary -- only by replicating across seeds; a stress test makes the stakes external: a single-run improvement of the size this field publishes would have the wrong sign roughly one-third of the time in a comparable regime. We release a library (cua_reliability) for routine k-seed reporting. The apparatus is, to our knowledge, the first multimodal segment-aggregated on-policy self-distillation (SA-OPSD) update on a real 35B CUA policy.
- Abstract(参考訳): エージェントコンピュータ利用RLは単一のランで報告され、それらの数値は誤解を招く。
5つのオーラルグレード環境にわたる35Bコンピュータ利用エージェント(CUA)のバリデーション誘導修復を用いて、修正されたポリシーの成功率は、上流の分散(crossed data-draw $\times$ seed grid, bootstrap CIs)による分散成分分解(crossed data-draw $\times$ seed grid, bootstrap CIs)によって支配される。
実行から実行までの分散はバイモーダルである(Hartigan dip $p=0.07$, $k=10$)。
2つの研究結果が得られた。
1つの固定トークンが確実にインストールされる(done-detection $0.97\pm0.06$)のに対して、オープンな修正は部分的な -- 空間座標クリック(grounding $0.53\pm0.35$)と生成フィールドフィル(0.14\pm0.04$)のみである。
第2に、修正アクションがタスクの唯一の残りのブロッカである場合にのみ、フレームレベルの修復がタスク成功に転送される( LinkedIn 8/20 vs. base 0/15, Fisher $p=0.006$)。
私たちは、サンプル効率曲線と'グラウンド化'境界境界の2つの過剰な評価を、種子を複製してのみ取得しました。ストレステストは、そのステストを外部に公開します。このフィールドが発行するサイズを1ランで改善すれば、同等の体制で、ほぼ3分の1のタイミングで間違ったサインが得られます。
通常のk-seedレポートのためのライブラリ(cua_reliability)をリリースする。
この装置は、我々の知る限り、実際の35B CUAポリシーに関する最初のマルチモーダルセグメント集約型自己蒸留(SA-OPSD)更新である。
関連論文リスト
- When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference [26.45899223541571]
データ駆動学習による事前知識の融合は、データが不足している場合に魅力的だが、それが役に立つ、冗長な、あるいは有害な場合には、制御されたアカウントが言うことはない。
我々は、データ駆動型代替(SimCLR、SimSiam、DINO、ImageNet転送、拡張、学習教師)に対して、トレーニング期間中に$sim$2%のオーバーヘッドでのみ注入されたGaborターゲットのピン留めされたバンクである固定手作りの知識ソースをベンチマークした。
測定したトレーニング時間の組み合わせ全体において、3つの結果が再帰する(決定レベル融合が異なる)。
論文 参考訳(メタデータ) (2026-08-21T13:44:10Z) - Test-time reasoning effort and unauthorized tool use in language-model agents: a prespecified equivalence study [1.6787220460106658]
ツールコールを通じて実行される言語モデルエージェントは、各ロールが実行する操作を制限するアクセス制御ポリシーの下で動作します。
TRIO-20の14の確認シナリオにおいて, GPT-5.6内の推論の努力(低, 最大)が異なる。
840の軌道と2つのモデル層にまたがって、認可されていないツールコールは発生しなかった。
論文 参考訳(メタデータ) (2026-08-04T06:01:14Z) - When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation [20.672156968511587]
合成計算タスクとSWE-bench Mini修正タスクにおける3アームアブレーション(ベースライン/bash_only/code_only)の比較を行った。
結果として得られた4つの(登録、エージェント)細胞全体で、エージェントを1つの execute_code MCP ツールに制限することは、最も安価なツールリッチなライバルである -- あるいは統計的に結びついている -- よりも安価である。
唯一の例外はSWE-bench/Claudeである。
論文 参考訳(メタデータ) (2026-07-12T04:52:08Z) - From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents [56.31499185764872]
教師の長い軌道上の監督された微調整(SFT)は、オープンソフトウェアエンジニアリング(SWE)エージェントに調査と推論を浸透させる主要な方法である。
本稿では,P2T (Patches-to-Trajectories) を提案する。P2T (Patches-to-Trajectories) は,P2T (Patches-to-Trajectories) において,P2T (Patches-to-Trajectories) とP2T (Patches-to-Trajectories) の2つの最適化法である。
論文 参考訳(メタデータ) (2026-05-21T04:54:55Z) - MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents [0.0]
検索強化エージェントに対するメモリ中毒攻撃を,統合評価フレームワークを用いたStackelbergゲームとして定式化する。
ASR-R: 0.25〜1.00$) による攻撃成功度を4倍に向上させる。
私たちの主な貢献は、勾配結合に接地したキャリブレーションに基づく防御であるMEMSADである。
論文 参考訳(メタデータ) (2026-05-05T08:15:41Z) - Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols [51.56484100374058]
そこで本研究では,単一プロトコルステップを正確なマッチングタスクで監査するためのペアアウトカム計測インタフェースを提案する。
各インスタンスについて、インターフェースはベースラインの正当性ビットと後ステップの正当性ビットを記録する。
これらのレートは精度の変化を予測し、種、混合物、パイプライン間でテスト可能な再利用可能な経験的インターフェースを定義する。
論文 参考訳(メタデータ) (2026-04-20T13:25:40Z) - Hardware Validation of DAGI via a Modular "Ridge" Signature and High-Order Synergistic Information [0.0]
IBM Quantumハードウェア上でのDAGI(Directed Acyclic Graph Information)フレームワーク。
理想的な出力分布が低次元モジュラー多様体(リッジ)に制約される小さな制御された実験
キーリカバリはチャンスを超えた:ショット毎の精度0.1689(チャンス0.125,95% Wilson CI[0.1610, 0.1772])
これらの結果は、DAGIが非自明でハードウェアに耐性のある情報構造を検出し、定量化するという主張を支持する。
論文 参考訳(メタデータ) (2026-04-16T14:16:59Z) - Tight Convergence Rates for Online Distributed Linear Estimation with Adversarial Measurements [66.94250413799232]
分散パラメータ-サーバ-ワーカー設定における乱数ベクトル$X$の推定について検討する。
主な課題は、敵の計測と非同期である。
その結果, 分散線形推定におけるロバスト性, 識別性, 統計的効率の統一的有限時間評価が得られた。
論文 参考訳(メタデータ) (2026-04-07T11:45:55Z) - High-Dimensional Robust Mean Estimation with Untrusted Batches [38.14592862692954]
本研究では,N$ユーザによるデータのコントリビューションを行う協調環境での高次元平均推定について検討した。
例えば、$varepsilon$-fraction of users is completely adversarial, and the more good' users provide data from distributions that related to $P$ but deviate by a near parameter $$.
我々のアルゴリズムは、最小最大誤差率$O(sqrtvarepsilon/n + sqrtd/nN + sを達成する。
論文 参考訳(メタデータ) (2026-02-24T08:59:37Z) - INC: An Indirect Neural Corrector for Auto-Regressive Hybrid PDE Solvers [61.84396402100827]
本稿では,学習した補正を支配方程式に統合する間接ニューラルコレクタ(mathrmINC$)を提案する。
$mathrmINC$は、$t-1 + L$の順番でエラー増幅を減らし、$t$はタイムステップ、$L$はリプシッツ定数である。
大規模なベンチマークで$mathrmINC$をテストし、1Dカオスシステムから3D乱流まで、多くの異なる解法、神経バックボーン、テストケースをカバーした。
論文 参考訳(メタデータ) (2025-11-16T20:14:28Z) - Certifiably Robust Model Evaluation in Federated Learning under Meta-Distributional Shifts [8.700087812420687]
異なるネットワーク "B" 上でモデルの性能を保証する。
我々は、原則付きバニラDKWバウンダリが、同じ(ソース)ネットワーク内の未確認クライアント上で、モデルの真のパフォーマンスの認証を可能にする方法を示す。
論文 参考訳(メタデータ) (2024-10-26T18:45:15Z) - Scalable 3D Registration via Truncated Entry-wise Absolute Residuals [65.04922801371363]
3ドルの登録アプローチでは、1000万ドル(107ドル)以上のポイントペアを、99%以上のランダムなアウトレイアで処理することができる。
我々はこの手法をTEARと呼び、Trncated Entry-wise Absolute Residualsを演算するoutlier-robust損失を最小限にする。
論文 参考訳(メタデータ) (2024-04-01T04:43:39Z) - Beyond Invariance: Test-Time Label-Shift Adaptation for Distributions
with "Spurious" Correlations [44.99833362998488]
テスト時のデータ分散の変化は、予測モデルのパフォーマンスに有害な影響を及ぼす可能性がある。
本研究では,未ラベルサンプルに適用したEMを用いて,共同分布の$p(y, z)$の変化に適応するテストタイムラベルシフト補正を提案する。
論文 参考訳(メタデータ) (2022-11-28T18:52:33Z) - Variance-Aware Confidence Set: Variance-Dependent Bound for Linear
Bandits and Horizon-Free Bound for Linear Mixture MDP [76.94328400919836]
線形バンドイットと線形混合決定プロセス(mdp)に対する分散認識信頼セットの構築方法を示す。
線形バンドイットに対しては、$d を特徴次元とする$widetildeo(mathrmpoly(d)sqrt1 + sum_i=1ksigma_i2) が成り立つ。
線形混合 MDP に対し、$widetildeO(mathrmpoly(d)sqrtK)$ regret bound を得る。
論文 参考訳(メタデータ) (2021-01-29T18:57:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。