論文の概要: Are Latent Reasoning Models Easily Interpretable?
- arxiv url: http://arxiv.org/abs/2604.04902v1
- Date: Mon, 06 Apr 2026 17:50:06 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-07 15:49:19.320612
- Title: Are Latent Reasoning Models Easily Interpretable?
- Title(参考訳): 遅延推論モデルは容易に解釈可能か?
- Authors: Connor Dilgren, Sarah Wiegreffe,
- Abstract要約: 潜在推論モデル(LRM)は推論コストの低さから研究の関心を集めている。
LRMは自然言語では意味がないため、監視が難しい。
本稿では,2つの最先端のLEMを検証し,LRMの解釈可能性について検討する。
- 参考スコア(独自算出の注目度): 8.215015010040917
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Latent reasoning models (LRMs) have attracted significant research interest due to their low inference cost (relative to explicit reasoning models) and theoretical ability to explore multiple reasoning paths in parallel. However, these benefits come at the cost of reduced interpretability: LRMs are difficult to monitor because they do not reason in natural language. This paper presents an investigation into LRM interpretability by examining two state-of-the-art LRMs. First, we find that latent reasoning tokens are often unnecessary for LRMs' predictions; on logical reasoning datasets, LRMs can almost always produce the same final answers without using latent reasoning at all. This underutilization of reasoning tokens may partially explain why LRMs do not consistently outperform explicit reasoning methods and raises doubts about the stated role of these tokens in prior work. Second, we demonstrate that when latent reasoning tokens are necessary for performance, we can decode gold reasoning traces up to 65-93% of the time for correctly predicted instances. This suggests LRMs often implement the expected solution rather than an uninterpretable reasoning process. Finally, we present a method to decode a verified natural language reasoning trace from latent tokens without knowing a gold reasoning trace a priori, demonstrating that it is possible to find a verified trace for a majority of correct predictions but only a minority of incorrect predictions. Our findings highlight that current LRMs largely encode interpretable processes, and interpretability itself can be a signal of prediction correctness.
- Abstract(参考訳): 潜在推論モデル(LRM)は、推論コストの低さ(明示的推論モデルに関連する)と、複数の推論経路を並列に探索する理論的能力により、大きな研究関心を集めている。
しかし、これらの利点は解釈可能性の低減のコストが伴う: LRMは自然言語では意味をなさないため、監視が難しい。
本稿では,2つの最先端のLEMを検証し,LRMの解釈可能性について検討する。
まず、遅延推論トークンは、論理推論データセットでは、遅延推論を全く使わずに、ほぼ常に同じ最終回答を生成できる。
この推論トークンの非活用は、なぜLEMが明示的な推論手法を一貫して上回らないのかを部分的に説明し、これらのトークンが先行作業で果たす役割について疑念を提起する。
第2に、遅延推論トークンがパフォーマンスに必要である場合、正しく予測されたインスタンスの65~93%の時間で、金の推論トレースをデコードできることを実証する。
このことは、LEMは解釈不能な推論プロセスではなく、期待されたソリューションを実装することが多いことを示唆している。
最後に,金の推論トレースを事前に知ることなく,潜在トークンから痕跡を復号する手法を提案し,正しい予測の大多数に対して検証されたトレースを見つけることは可能であるが,誤予測の少数しか見つからないことを示した。
以上の結果から,現在のLEMは解釈可能なプロセスの大部分をコード化しており,解釈可能性自体が予測精度の信号である可能性が示唆された。
関連論文リスト
- On the Self-awareness of Large Reasoning Models' Capability Boundaries [46.74014595035246]
本稿では,Large Reasoning Models (LRM) が機能境界の自己認識性を持っているかを検討する。
ブラックボックスモデルでは、推論式は境界信号を明らかにし、解決不可能な問題に対する信頼軌道は加速するが、解決不可能な問題に対する収束不確実軌道は加速する。
ホワイトボックスモデルでは,最後の入力トークンの隠れ状態が境界情報を符号化し,解答可能かつ解答不能な問題を推論開始前に線形分離可能であることを示す。
論文 参考訳(メタデータ) (2025-09-29T12:40:47Z) - Stop Spinning Wheels: Mitigating LLM Overthinking via Mining Patterns for Early Reasoning Exit [114.83867400179354]
オーバーライドは、大きな言語モデル全体のパフォーマンスを低下させる可能性がある。
推論は, 探索段階の不足, 補償推論段階, 推論収束段階の3段階に分類される。
我々は,ルールに基づく軽量なしきい値設定戦略を開発し,推論精度を向上させる。
論文 参考訳(メタデータ) (2025-08-25T03:17:17Z) - On Reasoning Strength Planning in Large Reasoning Models [50.61816666920207]
我々は, LRM が, 世代前においても, アクティベーションにおける推論強度を事前に計画している証拠を見出した。
次に、LEMがモデルのアクティベーションに埋め込まれた方向ベクトルによって、この推論強度を符号化していることを明らかにする。
我々の研究は、LEMにおける推論の内部メカニズムに関する新たな洞察を提供し、それらの推論行動を制御するための実践的なツールを提供する。
論文 参考訳(メタデータ) (2025-06-10T02:55:13Z) - Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning [33.040747962183076]
大規模推論モデル(LRM)は複雑な問題解決において顕著な能力を示したが、その内部の推論機構はよく理解されていない。
特定の生成段階におけるMIは, LRMの推論過程において, 突然, 顕著な増加を示す。
次に、これらのシンキングトークンがLRMの推論性能に不可欠であるのに対して、他のトークンは最小限の影響しか与えないことを示す。
論文 参考訳(メタデータ) (2025-06-03T13:31:10Z) - LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models [52.03659714625452]
最近開発された大規模言語モデル (LLM) は、幅広い言語理解タスクにおいて非常によく機能することが示されている。
しかし、それらは自然言語に対して本当に「理性」があるのだろうか?
この疑問は研究の注目を集めており、コモンセンス、数値、定性的など多くの推論技術が研究されている。
論文 参考訳(メタデータ) (2024-04-23T21:08:49Z) - Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation [110.71955853831707]
我々は、LMを、事前学習時に見られる間接的推論経路を集約することで、新たな結論を導出すると考えている。
我々は、推論経路を知識/推論グラフ上のランダムウォークパスとして定式化する。
複数のKGおよびCoTデータセットの実験と分析により、ランダムウォークパスに対するトレーニングの効果が明らかにされた。
論文 参考訳(メタデータ) (2024-02-05T18:25:51Z) - Self-Contradictory Reasoning Evaluation and Detection [31.452161594896978]
本稿では,自己矛盾推論(Self-Contra)について考察する。
LLMは文脈情報理解や常識を含むタスクの推論において矛盾することが多い。
GPT-4は52.2%のF1スコアで自己コントラを検出できる。
論文 参考訳(メタデータ) (2023-11-16T06:22:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。