論文の概要: The Delegation Danger Band: Why Mid-Capability Sub-Agents Over-Trust Inherited Stale State
- arxiv url: http://arxiv.org/abs/2610.00041v1
- Date: Thu, 03 Sep 2026 03:05:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-05 14:48:16.326705
- Title: The Delegation Danger Band: Why Mid-Capability Sub-Agents Over-Trust Inherited Stale State
- Title(参考訳): デリゲーション・ディガー・バンド:なぜ中間能力のサブエージェントが過度にトラストされたステールステートを継承したのか
- Abstract要約: C_m$ はクリーンなfork-fresh の精度を示す。
すべてのタスクはベースエビデンスから解決可能であるため、パフォーマンス損失は古い状態に依存する可能性がある。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Agent frameworks increasingly delegate work by forking sub-agents; a common default makes the child inherit the parent's full working context. We measure how the effect of inherited state changes with capability, where $C_m$ denotes clean fork-fresh accuracy. We compare 3 inheritance policies: Reset (fork fresh: base evidence only), Selective (curated handoff: + the useful prior conclusion), and Full (implicit fork: + the useful conclusion and $d$ copies of a superseded conclusion) over a same-family ladder (Qwen3 0.6/1.7/4/8B) on a frozen, closed-set, action-scored benchmark. Every task is solvable from the base evidence, so performance loss can be attributed to reliance on stale state. (1) Deference to superseded state falls sharply with measured capability $C_m$ (the slope's confidence interval, CI, excludes zero on every family) across 2 synthetic primitives plus MuSiQue and HotpotQA. (2) On the Qwen3 synthetic ladder, net inheritance harm follows a nonmonotone pattern: a mid-capability model (Qwen3-1.7B) is a statistically significant local minimum of net harm, falling below its fork-fresh baseline ($Δ(32)=-0.19$ [-0.25, -0.12]) and both neighbors, while the weakest model stays near-neutral and the strongest models stay robust. We call this harmful capability range a danger band. A within-model counting-difficulty sweep shows that the effect depends on model class even at matched $C_m$, and a live parent-to-child fork reproduces the mid-model harm. (3) Curated Selective handoff improves average accuracy over Full on all 3 datasets, largest at the in-band model, while the fixed-threshold capability router fails on the other datasets; a transferable router would need to predict the balance between reuse benefit and stale-context penalty. The benchmark is frozen and version-hashed.
- Abstract(参考訳): エージェントフレームワークは、サブエージェントをフォークすることで作業を委譲する傾向にある。
C_m$ はクリーンなfork-fresh の精度を示す。
Reset (fork fresh: base evidence only), Selective (curated handoff: + the useful pre conclusion), Full (implicit fork: + the useful conclusion and $d$ copy of a supersed conclusion) over a same-layer ladder (Qwen3 0.6/1.7/4/8B) on a frozen, closed-set, action-scored benchmark。
すべてのタスクはベースエビデンスから解決可能であるため、パフォーマンス損失は古い状態に依存する可能性がある。
1) 2つの合成プリミティブに加えて MuSiQue と HotpotQA の合計値である $C_m$ (傾きの信頼区間である CI は、各ファミリーのゼロを除外する) を伴って、過渡状態への参照が急激に低下する。
2) Qwen3合成ラグでは, 純継承害は非モノトンパターンに従う: 中間能力モデル (Qwen3-1.7B) は, 統計的に有意な局所的純害の最小値であり, フォーク・フレッシュベースライン(Δ(32)=-0.19$ [-0.25, -0.12]) と近隣の双方に降り注ぐ。
私たちはこの有害な能力は危険帯だと言っています。
in-model counting-difficulty sweepでは、一致した$C_m$でもモデルクラスに依存することが示され、生の親子フォークは中間モデルの害を再現する。
(3) キュレートされた選択ハンドオフは、バンド内モデルで最大となる3つのデータセットすべてに対して平均精度を向上する一方、固定閾値機能ルータは他のデータセットではフェールする。
ベンチマークは凍結され、バージョン管理される。
関連論文リスト
- SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models [53.86918766240095]
Delta-Ruleリカレントモデルは固定サイズの状態を維持しており、$O(1)$の推論メモリが可能であるが、極端なコンテキスト外挿では不安定になる可能性がある。
我々は,チャンク内並列構造を保ちながら,チャンク境界における適応$tanh$圧縮を適用したtextbfState Anomaly Neutralization (SANE)を提案する。
論文 参考訳(メタデータ) (2026-08-23T10:41:07Z) - Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk [0.0]
クレジットスコアリングは、決定ロジックをパラメータから読み取ることができないモデルに依存している。
共通する提案は、言語モデルとのギャップを埋める: 特徴属性を計算し、それらを LLM に渡し、合理的に記述させる。
このようなシステムをエンド・ツー・エンドに構築し、約束の後半が成立するかどうかをテストします。
論文 参考訳(メタデータ) (2026-08-08T13:22:14Z) - $f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data [3.294420397461204]
GFlowNetsと変分推論では、ターゲットとモデルログの確率の平均二乗誤差は、生成モデルのトレーニングに有効な、低分散、サロゲート損失である。
この損失は、$f$-divergences(英語版)のファミリー全体に拡張可能であることが示され、その結果、政治上の勾配が対応する$f$-divergence(英語版)のファミリーとなるが、同じグローバルな最小化のオフポリティ(英語版)を維持している。
この同値性により、対応する$f$-divergenceの特性を継承する幅広い生成モデルのクラスをチューニングするための新しい代理損失関数を設計できる。
論文 参考訳(メタデータ) (2026-05-14T21:02:07Z) - Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents [0.0]
自律的なAIエージェントは、完全に認証されたままで、振る舞いのドリフト、敵の適応、決定パターンのシフトによって、コードの変更なしに、安全が保たれる。
エージェントの管理は、未観測のリスクに対する限界を見積もることを減らす。
textbfRiskGateはこのフレームワークを、専用の統計推定器(KL分散、セグメント-vs-rest $z$-tests、シーケンシャルパターンマッチング)、フェイルセーフなモノトニックパイプライン、クローズドループオートパイロットでインスタンス化する。
論文 参考訳(メタデータ) (2026-04-27T16:46:15Z) - TRACE: Theoretical Risk Attribution under Covariate-shift Effects [4.211510706776732]
ソースをトレーニングしたモデル$Q$を、シフトしたデータに基づいてトレーニングされたモデル$tildeQ$に置き換えると、ソースドメインのパフォーマンスは予測不能に変化する可能性がある。
TRACEは$|R|$を解釈可能な上界に分解するフレームワークである。
論文 参考訳(メタデータ) (2026-02-11T07:22:33Z) - Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers [90.50039419576807]
RLVR(Reinforcement Learning with Verifiable Rewards)は、人為的なラベル付けを避けるために、自動検証に対するポリシーを訓練する。
認証ハッキングの脆弱性を軽減するため、多くのRLVRシステムはトレーニング中にバイナリ$0,1$の報酬を破棄する。
この選択にはコストがかかる:textitfalse negatives(正しい回答、FNを拒絶)とtextitfalse positives(間違った回答、FPを受け入れる)を導入する。
論文 参考訳(メタデータ) (2025-10-01T13:56:44Z) - Certifiably Robust Model Evaluation in Federated Learning under Meta-Distributional Shifts [8.700087812420687]
異なるネットワーク "B" 上でモデルの性能を保証する。
我々は、原則付きバニラDKWバウンダリが、同じ(ソース)ネットワーク内の未確認クライアント上で、モデルの真のパフォーマンスの認証を可能にする方法を示す。
論文 参考訳(メタデータ) (2024-10-26T18:45:15Z) - Certifiably Robust Interpretation via Renyi Differential Privacy [77.04377192920741]
我々はRenyi差分プライバシー(RDP)の新しい視点から解釈堅牢性の問題を研究する。
まず、証明可能で証明可能なトップ$k$ロバスト性を提供する。
第二に、提案手法は既存の手法よりも実験的堅牢性を$sim10%$で提供する。
第3に,ロバスト性と計算効率のトレードオフを円滑に行うことができる。
論文 参考訳(メタデータ) (2021-07-04T06:58:01Z) - The Curse of Performance Instability in Analysis Datasets: Consequences,
Source, and Suggestions [93.62888099134028]
自然言語推論(NLI)および読み込み(RC)解析/ストレスセットにおける最先端モデルの性能は極めて不安定であることがわかった。
このことは、(1)不安定さがこれらの分析セットに基づいて引き出された結論の信頼性にどのように影響するかという3つの疑問を提起する。
不安定の原因に関する理論的説明と実証的証拠の両方を提示する。
論文 参考訳(メタデータ) (2020-04-28T15:41:12Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。