論文の概要: Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification
- arxiv url: http://arxiv.org/abs/2609.06445v1
- Date: Sun, 06 Sep 2026 07:36:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-10 19:44:08.252022
- Title: Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification
- Title(参考訳): エージェント決定の因果属性:推定器、結合、トレーサビリティの仕様
- Abstract要約: 推定器のフレームワークと、それが失敗する条件を与えます。
限界推定の下では、因果不活性ステップは、植えられた鎖の全ての実行において決定的なステップと同一の合計効果を有する。
コンテキストが分岐すると、直接効果を推定できる結合を導出します。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: A provider of a high-risk AI system must keep records that make a decision traceable, and for agentic systems it has not been established what those records must contain for post-hoc causal attribution to be possible. We give the estimator framework and then the conditions under which it fails. We separate the marginal total effect that prior work measures from a common-random-numbers total effect that isolates a step's own contribution, add the natural direct effect under a pinned downstream, and check the estimators against hand derivations. Both estimands then fail, in the same direction. Under the marginal estimand a causally inert step has the identical total effect to the decisive one on every run of our planted chain, an algebraic identity and not a coincidence at one draw. Under common random numbers the decisive step returns exactly zero on the runs where the executing step flips, about one in ten, while its direct effect there is 0.25 and it demonstrably acts; an exact zero does not certify that a step did nothing, and we put that here rather than in the limitations. We derive the coupling that keeps the direct effect estimable once contexts diverge, with a closed form for its degradation, and show that the mediated share on which a natural ranking is built is not a share under suppression: where the direct and mediated paths oppose, it exceeds one and ranks a suppressed component above a pure mediator. We publish the discrepancy experiment's pre-registration rather than a result, because the live pipeline it requires was not available in the study window. We contribute the traceability specification such a filing would need, against a gap the Act's calendar opens: Article 86's right to an explanation has applied since 2 August 2026, while the Article 12 logging and Annex IV documentation that could evidence one were deferred to 2 December 2027 by Regulation (EU) 2026/1744.
- Abstract(参考訳): リスクの高いAIシステムのプロバイダは、意思決定をトレース可能なレコードを保持しなければならない。
推定器のフレームワークと、それが失敗する条件を与えます。
本研究は,従来の作業方法と,ステップのコントリビューションを分離する共通ランダム数のトータルエフェクトとを分離し,ピン留めされた下流で自然の直接効果を付加し,手動の導出に対する推定値をチェックすることによる限界的トータルエフェクトを分離する。
両方の推定値は同じ方向に失敗する。
限界推定の下では、因果不活性なステップは、植え付け鎖のすべての実行において決定的なステップと同一の合計効果を持ち、代数的同一性は1つの引き分けにおいて偶然ではない。
一般的な乱数の下では、決定的なステップは、実行中のステップがフリップするランで正確にゼロを返すが、その直接的な効果は0.25であり、明らかに作用する。
また, 直接経路と媒介経路が反対の場合には, 抑制された成分を純粋メディエータの上にランク付けし, 自然なランク付けを行う媒介共有が抑制対象のシェアではないことを示す。
研究ウィンドウでは必要となるライブパイプラインが利用できなかったため,結果ではなく,不一致実験の事前登録を公開しました。
第86条の説明に対する権利は2026年8月2日から適用されており、第12条のロギング及びAnnex IVの文書は2027年12月2日にEU規則(EU)2026/1744により延期された。
関連論文リスト
- The Model Proposes, the Code Disposes: A Pre-Registered Ablation of a Verifier-and-Acceptance Stage in an LLM-Orchestrated Offensive-Security Agent [0.0]
決定的コードによって検証が強制されるモデル検証装置である検証受理段階が、LCMが主導する攻撃セキュリティエージェントが報告するものを変更するかどうかを評価する。
本報告では,15回の探索パイロット,20回の予備登録検定アブレーション,および2回の故意に脆弱な検査対象を40回実施した2×2因子分析を行った。
検証者はシステムが提供するものを変更し、決定論的コードは強制と監査性を提供します。
論文 参考訳(メタデータ) (2026-09-14T17:13:05Z) - Principled Detection of Coordinated Manipulation from Aggregate Distortion and Account Reuse [10.023638697449062]
調整可能なアカウントは、評価、ランキング、エンゲージメントを共同で歪めることができる。
我々は,文脈レベルの結果分布の歪みを主証拠として扱うアグリゲートファーストエビデンス層を導入する。
制御された回転実験とペアの対実的介入によるメカニズムを過去のAmazonレビューストリームで評価した。
論文 参考訳(メタデータ) (2026-09-11T18:16:42Z) - Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay [0.6768558752130311]
単一エージェント環境において,実行されたリプレイからポリシー条件の真偽を検査する。
段階的な信用シグナルは、限界整合シャッフル制御以上の信頼度の高いインクリメンタル忠実度を示すものではない。
信頼のみのルータは、チャンスレベルにおいて重要なステップを回復するが、1ターンあたりの審査コストを13.1%削減する。
論文 参考訳(メタデータ) (2026-08-20T08:04:00Z) - Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce [39.937871162878785]
我々は、虚偽の事実主張、操作、共謀、脅迫を含む電子メールとして、音声行為の誤認識を運用する。
メールの12.6%はミスアライメントと表示されており、20行中、74.7%は個別のエージェントランで表示されている。
高い能力モデルが弱い候補を差分に利用しているという証拠は見つからない。
論文 参考訳(メタデータ) (2026-08-14T18:56:12Z) - When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide [0.0]
5つのデータセットと2つの既知のエフェクトスイープにまたがる6つの推定器をベンチマークし、非シミュレートされたペア参照に対してそのメカニズムを検証する。
ターゲット自身のスコアからロガーをシャープすると、テスト対象範囲をわずかにオーバーラップし、アクションレベルの不一致が崩壊する。
結果のニュアンスのみを適合させることは、再利用バイアスをその場に残し、それを悪化させます。
論文 参考訳(メタデータ) (2026-08-12T18:10:10Z) - Orientation, not magnitude: the causal structure of task-vector interference in merged language models [0.0]
タスク算術はそうでなければ機能し、フィールドはマグニチュードで原因を診断する。
層状フラックスの正確な分解は、既存の断続輸送の増幅によって支配されていることを示している。
ナイーブのbfloat16生成の「ユニバーサリティ」は、量子化粗さであることが判明した。
論文 参考訳(メタデータ) (2026-08-12T08:40:24Z) - A Control Theory of Predictability in Latent World Models [39.22646915328561]
現在のプラクティスでは、トレーニングとモデル選択の目的として、ホールトアウトデータに対する単一または複数ステップのロールアウト損失である予測エラーを採用しています。
この仮定は構造上の理由から信頼できないことを示す。
プランナーはトレーニング分布についてモデルに問い合わせるのではなく、その候補となるアクションが到達し、一般にデータ多様体を残します。
論文 参考訳(メタデータ) (2026-07-11T15:42:26Z) - Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures [0.0]
Causal Agent Replay (CAR) は、エージェントが構造的因果モデルとして実行されることをモデル化する。
ステップにダブルオペレーションを適用し、同じポリシーの下で軌道を再実行します。
CARはオープンソースで、ホストまたはフリーのローカルモデルで動作する。
論文 参考訳(メタデータ) (2026-06-06T17:44:23Z) - Causal Understanding by LLMs: The Role of Uncertainty [43.87879175532034]
近年の論文では、LLMは因果関係分類においてほぼランダムな精度を達成している。
因果的事例への事前曝露が因果的理解を改善するか否かを検討する。
論文 参考訳(メタデータ) (2025-09-24T13:06:35Z) - Nonparametric Identifiability of Causal Representations from Unknown
Interventions [63.1354734978244]
本研究では, 因果表現学習, 潜伏因果変数を推定するタスク, およびそれらの変数の混合から因果関係を考察する。
我々のゴールは、根底にある真理潜入者とその因果グラフの両方を、介入データから解決不可能なあいまいさの集合まで識別することである。
論文 参考訳(メタデータ) (2023-06-01T10:51:58Z) - Treatment Effect Risk: Bounds and Inference [58.442274475425144]
平均的な治療効果は社会福祉の変化を測定するため、たとえ肯定的であっても、人口の約10%に悪影響を及ぼすリスクがある。
本稿では,ICT分布のリスク条件値(CVaR)として定式化されたこの重要なリスク尺度をどう評価するかを検討する。
いくつかの境界は、複素CATE関数を単一の計量に要約したものと解釈することもでき、有界であることとは無関係に興味を持つ。
論文 参考訳(メタデータ) (2022-01-15T17:21:26Z) - Nested Counterfactual Identification from Arbitrary Surrogate
Experiments [95.48089725859298]
観測と実験の任意の組み合わせからネスト反事実の同定について検討した。
具体的には、任意のネストされた反事実を非ネストされたものへ写像できる反ファクト的非ネスト定理(英語版)(CUT)を証明する。
論文 参考訳(メタデータ) (2021-07-07T12:51:04Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。