論文の概要: The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
- arxiv url: http://arxiv.org/abs/2606.00376v1
- Date: Fri, 29 May 2026 21:35:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-02 21:34:28.369626
- Title: The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
- Title(参考訳): 決定論的水平線:拡張推論障害とツールデリゲーションが必須になる
- Authors: Dongxin Guo, Jikun Wu, Siu Ming Yiu,
- Abstract要約: 拡張連鎖推論は決定論的状態追跡タスクのパフォーマンスを低下させる。
ツール統合推論がニューラル・チェーン・オブ・シントを一貫して上回ることを示す。
本研究は, エージェントシステムにおいて, 純粋なニューラル推論がハイブリッドアプローチにいつ適用されるべきかについて, 基本的ガイダンスを提供する。
- 参考スコア(独自算出の注目度): 13.891522069967507
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not due to preference biases, but limits rooted in the information-theoretic capacity of decoder-only attention. We establish: (1) an Attention Bottleneck Theorem with a complementary achievability construction, bounding state-tracking capacity as $O(H \cdot \log(L/H) \cdot \sqrt{d_h})$; (2) a context-dependent error model yielding super-exponential accuracy decay; (3) the State-Space Jaccard metric distinguishing capability from preference failures; (4) a Deterministic Horizon $d^* \in [19, 31]$ beyond which tool delegation becomes necessary. Across 12 models and 8 task domains (including SWE-Bench, WebArena, and SQL-Multi), tool-integrated reasoning consistently outperforms neural chain-of-thought; on the primary model suite it reaches 86-94% accuracy versus 24-42% for neural chain-of-thought. Fine-tuning on optimal-length traces yields $<$5% improvement, confirming an architectural ceiling, and high cross-model correlation ($r = 0.81$-$0.91$) indicates these failures are architectural rather than training-specific. Our results provide principled guidance for when pure neural reasoning should yield to hybrid approaches in agentic systems.
- Abstract(参考訳): 拡張チェーン・オブ・ソート推論は、優先バイアスによるものではなく、決定論的状態追跡タスクのパフォーマンスを低下させることができるが、デコーダのみの注意力の情報理論能力に根ざした制限がある。
Atention Bottleneck Theorem with a complementary achievability construction, bounding state-tracking capacity as $O(H \cdot \log(L/H) \cdot \sqrt{d_h})$; (2) a context-dependent error model yielding super-exponential accuracy decay; (3) State-Space Jaccard metric metric distinguishing capabilities from preference failures; (4) a Deterministic Horizon $d^* \in [19, 31]$。
12のモデルと8つのタスクドメイン(SWE-Bench、WebArena、SQL-Multiを含む)で、ツール統合推論は一貫してニューラルチェーン・オブ・シントを上回っている。
最適長トレースの微調整は、アーキテクチャの天井を確認し、高いクロスモデル相関(r = 0.81$-$0.91$)は、これらの障害がトレーニング固有のものではなくアーキテクチャであることを示している。
本研究は, エージェントシステムにおいて, 純粋なニューラル推論がハイブリッドアプローチにいつ適用されるべきかについて, 基本的ガイダンスを提供する。
関連論文リスト
- Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning [59.74608632210439]
そこで本研究では,ツール使用の自然な動作を,ツールなし推論能力を犠牲にすることなく,強力な思考モデルに注入する方法を示す。
提案手法は,オープンソースモデル間のベンチマークにおいて,最先端のパフォーマンスを実現するモデルを生成する。
論文 参考訳(メタデータ) (2026-05-07T14:23:21Z) - Evaluating Small Language Models for Front-Door Routing: A Harmonized Benchmark and Synthetic-Traffic Experiment [0.05586191108738562]
小型言語モデル(SLM)は、サブ秒、ゼロマージナルコスト、セルフホストタスクの分類に十分な推論能力を持つ。
Study 1はPhi-3.5-mini、Qwen2.5-1.5B、Qwen-2.5-3Bを同一のAzure T4ハードウェア、サービススタック、量子化、固定60ケースコーパスで同期したオフラインベンチマークである。
研究2は、合成トラフィック下で事前登録された4本腕ランダム化実験であり、有効サンプルサイズは腕あたり60ケースである。
論文 参考訳(メタデータ) (2026-03-26T15:57:46Z) - Differentiable Symbolic Planning: A Neural Architecture for Constraint Reasoning with Learned Feasibility [0.0]
微分可能シンボリックプランニング(DSP)は、離散的シンボリック推論を行うニューラルネットワークである。
DSPをUniversal Cognitive Kernel(UCK)に統合し、グラフ注意と反復的制約伝搬を組み合わせた。
UCK+DSPは4倍の精度で計画の精度を97.4%向上させる。
論文 参考訳(メタデータ) (2026-02-19T03:38:03Z) - ReLE: A Scalable System and Structured Benchmark for Diagnosing Capability Anisotropy in Chinese LLMs [37.23311145049677]
本稿では,機能異方性(Capability Anisotropy)を診断するためのスケーラブルなシステムであるReLEを提案する。
我々は,207,843サンプルからなる領域$times$ Capability SymbolicMatrixの304モデルを評価した。
論文 参考訳(メタデータ) (2026-01-24T09:57:59Z) - Towards a Science of Scaling Agent Systems [79.64446272302287]
エージェント評価の定義を定式化し,エージェント量,コーディネーション構造,モデル,タスク特性の相互作用として,スケーリング法則を特徴付ける。
協調指標を用いて予測モデルを導出し,R2=0をクロスバリデーションし,未知のタスク領域の予測を可能にする。
ツールコーディネーショントレードオフ: 固定的な計算予算の下では, ツールヘビータスクはマルチエージェントのオーバーヘッドから不均衡に悩まされ, 2) 能力飽和: 調整が減少または負のリターンを, 単一エージェントのベースラインが45%を超えると達成できる。
論文 参考訳(メタデータ) (2025-12-09T06:52:21Z) - vAttention: Verified Sparse Attention [100.98210818821688]
vAttentionは、ユーザが指定した$(epsilon, delta)$の近似精度保証(thus, confirmed)を備えた実用的なスパースアテンションメカニズムである。
vAttentionはデータセット間のスパースアテンションの質を大幅に改善することを示す。
モデルの品質を損なうことなく高速なデコードを実現するために、推論シナリオにデプロイすることができる。
論文 参考訳(メタデータ) (2025-10-07T08:46:08Z) - Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning [53.45095336430027]
暗黙的な検索と構造化された協調を組み合わせた統合フレームワークを開発する。
Humanity's Last Exam (HLE) Bio/Chem Goldでは,48.3%の精度を実現している。
SuperGPQAとTRQAの結果はドメイン間の堅牢性を確認した。
論文 参考訳(メタデータ) (2025-09-25T14:05:55Z) - Understanding, Predicting and Better Resolving Q-Value Divergence in
Offline-RL [86.0987896274354]
まず、オフラインRLにおけるQ値推定のばらつきの主な原因として、基本パターン、自己励起を同定する。
そこで本研究では,Q-network の学習における進化特性を測定するために,SEEM(Self-Excite Eigen Value Measure)尺度を提案する。
われわれの理論では、訓練が早期に発散するかどうかを確実に決定できる。
論文 参考訳(メタデータ) (2023-10-06T17:57:44Z) - NPBDREG: A Non-parametric Bayesian Deep-Learning Based Approach for
Diffeomorphic Brain MRI Registration [0.0]
NPBDREGは、教師なしの変形可能な画像登録のための非最適フレームワークである。
それは理論上よくパラメトリックで計算的に効率的な方法で、改善された不確実性推定と信頼度測定を提供する。
PrVXMに比べて登録精度は若干改善されている。
論文 参考訳(メタデータ) (2021-08-15T16:00:06Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。