論文の概要: From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought
- arxiv url: http://arxiv.org/abs/2609.25366v1
- Date: Mon, 21 Sep 2026 20:06:06 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-23 18:04:04.095608
- Title: From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought
- Title(参考訳): 定型化からロードバアリングへ: 作業難易度は, 結束の因果的役割を形作る
- Abstract要約: CoT(Chain-of- Thought)モニタリングは、記述された推論が答えを因果的に制約する場合のみ意味を持つ。
本稿では,連続型因果検定,アブレーション・パッチによる介入,連鎖の切断,崩壊した接頭辞からの継続を強制する手法を提案する。
これは、CoTが最後に答えるためにどのようにロードされるかを測定するもので、機械的忠実性とは異なる振る舞いの概念である。
- 参考スコア(独自算出の注目度): 2.122752621320654
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Chain-of-thought (CoT) monitoring is only meaningful if written reasoning causally constrains the answer. We introduce continuation-based causal testing, an ablation-patch intervention that perturbs one reasoning step, truncates the chain, and forces the model to continue from the corrupted prefix. It measures how load-bearing a CoT is for the final answer, a behavioral notion distinct from mechanistic faithfulness. Across Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness tracks model-relative task difficulty: on easy tasks models silently bypass their own reasoning; on hard tasks they follow corrupted steps and propagate errors. A matched 2x2 analysis shows task difficulty dominates perturbation type: error propagation rises 16x from GSM8K to BBH multistep arithmetic, and a variance partition over 28,584 continuations attributes 98.8% of explained deviance to task difficulty versus 0.8% to perturbation type. Reasoning-specific RL suppresses error propagation and compresses the gradient. A four-variant judge-sensitivity analysis and blind two-annotator study (n=500) show the error-propagation vs. non-propagation label is invariant to judge prompt, with perfect inter-annotator agreement (Cohen's kappa = 1.00). This gradient creates a structural problem for CoT-based oversight and AI safety monitoring: where the trace is easy to read it carries little signal, and where it matters errors propagate before a monitor can intervene. Linear probes on hidden states separate silent bypass, self-correction, and error propagation, but additive activation steering provides limited causal control, flipping only about 25% of error-propagation cases at best. Behavioral mode is readable but not reliably controllable.
- Abstract(参考訳): CoT(Chain-of- Thought)モニタリングは、記述された推論が答えを因果的に制約する場合にのみ意味を持つ。
本稿では,連続型因果検定,アブレーション・パッチによる介入,連鎖の切断,崩壊した接頭辞からの継続を強制する手法を提案する。
これは、CoTが最後に答えるためにどのようにロードされるかを測定するもので、機械的忠実性とは異なる振る舞いの概念である。
Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness track model-relative task difficulty: on easy task models on quietly bypass their reasoning; on hard task they follow corrupted steps and propagate error。
一致した2x2分析は、タスクの難易度が摂動型を支配していることを示している: エラーの伝播は GSM8K から BBH 多段階算術へ 16x 上昇し、28,584 の継続特性上の分散分割は、タスクの難易度に98.8%、摂動型は 0.8% である。
推論固有のRLはエラーの伝播を抑制し、勾配を圧縮する。
4変量判定感度解析とブラインド2アノテーション研究(n=500)は、エラープロパゲーション対非プロパゲーションラベルは、完全なアノテータ間合意(Cohen's kappa = 1.00)で、プロンプトを判断するために不変であることを示している。
この勾配は、CoTベースの監視とAIの安全監視のための構造的な問題を生み出します。
隠れ状態上の線形プローブはサイレントバイパス、自己補正、エラーの伝播を分離するが、加算活性化ステアリングは因果制御を制限し、エラープロパゲーションのケースの25%しか反転しない。
動作モードは読みやすいが、確実に制御できない。
関連論文リスト
- Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability [51.56484100374058]
機械論的解釈可能性プリミティブに対する最初の識別可能性定理を証明した。
スペクトルは、正当性分解ではなく、記述されたエラーバーを持つ識別可能なモデル固有の指紋である。
論文 参考訳(メタデータ) (2026-08-10T19:42:01Z) - Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning [109.33717747526583]
大規模言語モデルは、計算コストが高くても意味論的に無効な推論を、能力を超えたタスクで生成する。
textbfCaRL(textbfCapability-textbfaligned textbfReinforcement textbfL)を導入し、振る舞いを機能境界に合わせる。
実験は、タスクの難易度を超えて性能を保ちながら、無駄な推論を大幅に削減することを示した。
論文 参考訳(メタデータ) (2026-07-31T09:30:33Z) - When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents [0.0]
長い答えのLSMエージェントは静かに失敗する可能性があり、彼らは証拠を早期に読み上げ、残りの期間をその証拠を守るのに費やした。
我々は、表現的コミットメントを、固定された推論ステップにおいて、クロスランな隠れ状態収束として定義する。
ランタイムモニタは、AUROCの隠れ状態から0.97までの不整合軌道を検出する(より厳密なスプリットの下で0.85-0.88)。
論文 参考訳(メタデータ) (2026-06-22T07:13:13Z) - Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought [11.955186033088351]
推論モデルにおける行動連鎖(CoT)の証拠を提供する。
アクティベーションプロービング、早期強制応答、および2つの大きなモデルにわたるCoTモニターを比較した。
難解なマルチホップGPQA-ダイアモンド問題における真の推論とは対照的である。
論文 参考訳(メタデータ) (2026-03-05T18:55:16Z) - Information Fidelity in Tool-Using LLM Agents: A Martingale Analysis of the Model Context Protocol [69.11739400975445]
モデルコンテキストプロトコル(MCP)エージェントにおけるエラー蓄積を解析するための最初の理論的枠組みを紹介する。
累積歪みが線形成長と高確率偏差を$O(sqrtT)$で表すことを示す。
主な発見は、意味重み付けは歪みを80%減らし、周期的再接地は、エラー制御の約9ステップごとに十分である。
論文 参考訳(メタデータ) (2026-02-10T21:08:53Z) - On the Paradoxical Interference between Instruction-Following and Task Solving [50.75960598434753]
次の命令は、大規模言語モデル(LLM)を、タスクの実行方法に関する明示的な制約を指定することで、人間の意図と整合させることを目的としている。
我々は,LLMのタスク解決能力にパラドックス的に干渉する命令に従うという,直感に反する現象を明らかにした。
本稿では,タスク解決に追従する命令の干渉を定量化する指標として,SUSTAINSCOREを提案する。
論文 参考訳(メタデータ) (2026-01-29T17:48:56Z) - Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems [20.846301581161978]
マルチエージェントシステムにおける障害帰属は、批判的だが未解決の課題である。
現在の手法では、これを長い会話ログ上のパターン認識タスクとして扱う。
A2P Scaffoldingは、パターン認識から構造化因果推論タスクへの障害帰属を変換する。
論文 参考訳(メタデータ) (2025-09-12T16:51:15Z) - Adapt in the Wild: Test-Time Entropy Minimization with Sharpness and Feature Regularization [85.50560211492898]
テスト時適応(TTA)は、テストデータが分散シフトが混在している場合、モデルの性能を改善または損なう可能性がある。
これはしばしば、既存のTTAメソッドが現実世界にデプロイされるのを防ぐ重要な障害である。
両面からTTAを安定化させるため,SARと呼ばれる鋭く信頼性の高いエントロピー最小化手法を提案する。
論文 参考訳(メタデータ) (2025-09-05T10:03:00Z) - CoT-Valve: Length-Compressible Chain-of-Thought Tuning [50.196317781229496]
我々はCoT-Valveと呼ばれる新しいチューニングと推論戦略を導入し、モデルが様々な長さの推論連鎖を生成できるようにする。
我々は,CoT-Valveがチェーンの制御性と圧縮性を実現し,プロンプトベース制御よりも優れた性能を示すことを示す。
論文 参考訳(メタデータ) (2025-02-13T18:52:36Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。