論文の概要: Is Code Better Than Language for Algorithmic Reasoning
- arxiv url: http://arxiv.org/abs/2606.15589v1
- Date: Sun, 14 Jun 2026 04:17:21 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-16 16:21:33.706625
- Title: Is Code Better Than Language for Algorithmic Reasoning
- Title(参考訳): アルゴリズム推論のための言語よりもコードの方が優れているか?
- Authors: Terry Tong, Yu Feng, Surbhi Goel, Dan Roth,
- Abstract要約: ツール拡張言語モデルでは、中間表現と実行機構の両方を変えるため、自然言語推論とコード実行パイプラインを比較することは困難である。
モデルは、その推論を実行可能なコードとして表現し、言語モデルは、そのコードを文脈でシミュレートして、回答を生成する。
中間介入は自然言語の推論と有意に異なるものではない(+0.15pp)。
- 参考スコア(独自算出の注目度): 58.86873062890482
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: For tool-augmented language models, comparing natural-language reasoning with code-execution pipelines is difficult because the comparison changes both the intermediate representation and the execution mechanism. We separate these factors with an intermediate intervention: the model expresses its reasoning as executable code, and the language model simulates that code in context to produce an answer. On a 40-task verifiable algorithmic benchmark, deterministic code execution outperforms natural-language reasoning by +31.6pp. We observe that the intermediate intervention is not meaningfully different from natural-language reasoning (+0.15pp). These results suggest that, in our evaluated setting, changing the intermediate representation alone does not explain the tool-use advantage, providing evidence for the performance gains requiring reliable external execution. We formalize this intuition with a simple statistical decision-theoretic model that characterizes when execution dominates end-to-end risk in our disentangled trace-generation/execution regime. We validate our theory using a reconstruction intervention that leverages a proxy language model to infer natural-language reasoning traces from code representations, recovering performance comparable to the original natural-language reasoning pipeline. All experiments are at https://github.com/TerryTong-Git/ToolProj.
- Abstract(参考訳): ツール拡張言語モデルでは、中間表現と実行機構の両方を変えるため、自然言語推論とコード実行パイプラインを比較することは困難である。
モデルは、その推論を実行可能なコードとして表現し、言語モデルは、そのコードを文脈でシミュレートして、回答を生成する。
40タスクの検証可能なアルゴリズムベンチマークでは、決定論的コードの実行は+31.6ppの自然言語推論よりも優れている。
中間介入は自然言語の推論(+0.15pp)と有意に異なるものではない。
これらの結果から,中間表現の変更だけではツール利用の利点を説明できないことが示唆され,信頼性の高い外部実行を必要とする性能向上の証拠となる。
我々はこの直観を単純な統計的決定理論モデルで定式化し、不整合なトレース生成/実行体制において、実行がエンドツーエンドのリスクを支配しているときに特徴付ける。
我々は、プロキシ言語モデルを利用して、コード表現から自然言語推論トレースを推論し、元の自然言語推論パイプラインに匹敵する性能を回復する再構成介入を用いて、我々の理論を検証する。
すべての実験はhttps://github.com/TerryTong-Git/ToolProj.comにある。
関連論文リスト
- Generating Verifiable CoT from Execution-Traces [6.634229408414094]
チェーン・オブ・ソート(Chain-of-Thought)のプロンプトは有望だが、現在の総合的なトレーニングデータは重大な弱点に悩まされている。
プログラム実行トレースにCoT生成を直接接地することで、この問題に対処する。
この実行基盤のアプローチは、プログラムが真に計算したものを反映するすべての推論ステップを保証する。
論文 参考訳(メタデータ) (2025-11-28T07:43:43Z) - Are Language Models Efficient Reasoners? A Perspective from Logic Programming [109.47572890883248]
現代言語モデル(LM)は、強い推論能力を示すが、標準的な評価は、人間のような推論の重要な側面である効率性を見越しながら、正確性を強調する。
本稿では、論理プログラミングのレンズを用いて、LM推論効率を評価するためのフレームワークを提案する。
論文 参考訳(メタデータ) (2025-10-29T15:30:31Z) - DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models [58.439517684779936]
本稿では,多種多様な文からなる自然文からなる古典論理ベンチマークDivLogicEvalを提案する。
また,より信頼性の高い評価を実現するために,大規模言語モデルに固有のバイアスやランダム性の影響を緩和する新たな評価指標を導入する。
論文 参考訳(メタデータ) (2025-09-19T04:40:46Z) - Learning to Disentangle Latent Reasoning Rules with Language VAEs: A Systematic Study [37.52166353495979]
本研究は,推論規則を言語モデル内に明示的に組み込んで記憶する方法について検討する。
本稿では,Transformer ベースの言語 VAE における推論規則を学習するための完全なパイプラインを提案する。
論文 参考訳(メタデータ) (2025-06-24T08:38:03Z) - Preventing Language Models From Hiding Their Reasoning [0.0]
大規模言語モデル(LLM)は、複雑な問題に対する答えを生成するための推論の中間ステップの恩恵を受けることが多い。
この研究では、推論の中間段階が不信である可能性のある1つの潜在的方法、すなわち符号化推論に焦点を当てる。
言語モデルは、ユーザが推論の中間ステップを理解せずに、符号化推論を利用してより高い性能を得るように訓練できることを示す。
論文 参考訳(メタデータ) (2023-10-27T22:02:29Z) - Rationale-Augmented Ensembles in Language Models [53.45015291520658]
我々は、数発のテキスト内学習のための合理化促進策を再考する。
我々は、出力空間における合理的サンプリングを、性能を確実に向上させるキーコンポーネントとして特定する。
有理拡張アンサンブルは既存のプロンプト手法よりも正確で解釈可能な結果が得られることを示す。
論文 参考訳(メタデータ) (2022-07-02T06:20:57Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。