論文の概要: Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
- arxiv url: http://arxiv.org/abs/2607.20520v1
- Date: Wed, 08 Jul 2026 17:21:28 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-27 00:46:13.215234
- Title: Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
- Title(参考訳): 数学的問題解決のための大規模言語モデルにおける実行可能推論制約下での表現ロバスト性
- Abstract要約: 本稿では,大規模言語モデル(LLM)における表現ロバスト性について検討する。
我々は、物語、記号、単語方程式の変種にまたがる非自明なフリップレートで、かなりの表現感度を見出した。
LLMの評価と展開において、表現は第一級インタフェース設計変数として扱われるべきである。
- 参考スコア(独自算出の注目度): 3.0618862102164996
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formulations as interchangeable and conflates reasoning errors with interface failures. This paper investigates representation robustness in LLM-based mathematical problem solving by systematically varying surface representations of the same underlying problems, including story problems, word-equations, symbolic equations, and isomorphic paraphrases. Using a curated dataset of mathematically equivalent problems, we evaluate five contemporary LLMs under a direct answer generation condition. We find substantial representational sensitivity: models frequently change correctness across equivalent formulations, with nontrivial flip rates across story, symbolic, and word-equation variants. We also observe systematic regressions under isomorphic reformulations, showing that even subtle paraphrase-level changes can degrade performance despite preserved mathematical structure. We then evaluate a code-augmented condition in which models externalize reasoning as executable Python code that is run locally for validation. This interface reveals strong latent reasoning capability in some models that perform poorly under direct prompting, but it does not uniformly improve robustness. Instead, failures shift across interaction layers, from opaque reasoning errors to protocol violations and execution failures. Even when executable reasoning succeeds, representation sensitivity often persists. Overall, our results show that reasoning scaffolds do not eliminate representational brittleness, but expose new tradeoffs among correctness, reliability, latency, and cost. We argue that representation should be treated as a first-class interface design variable in LLM evaluation and deployment, especially for AI-assisted problem-solving systems.
- Abstract(参考訳): 大規模言語モデル (LLM) は数学的な問題解決においてますます評価されているが、以前の研究は表現的に等価な定式化を交換可能なものとして扱い、インタフェースの故障を伴う推論エラーを混同する。
本稿では,物語問題,単語方程式,記号方程式,同型パラフレーズなど,同じ問題の表面表現を体系的に変化させることにより,LLMに基づく数学的問題解決における表現ロバスト性について検討する。
数学的に等価な問題のキュレートされたデータセットを用いて、5つの現代のLCMを直接解答生成条件下で評価する。
モデルはしばしば等価な定式化にまたがって正当性を変化させ、ストーリー、記号、単語方程式の変種を非自明なフリップレートで変更する。
また,同型再定式化による系統的回帰も観察し,微妙なパラフレーズレベルの変化であっても,保存された数学的構造にもかかわらず性能が低下することを示した。
次に、モデルが推論を検証のためにローカルに実行される実行可能なPythonコードとして外部化する、コード拡張条件を評価する。
このインタフェースは、直接的プロンプト下では性能が良くないモデルでは強い潜伏推論能力を示すが、ロバスト性は均一に改善しない。
代わりに、障害は、不透明な推論エラーからプロトコル違反、実行障害へと、インタラクション層を横断します。
実行可能な推論が成功したとしても、表現感度は持続する。
総じて,足場の推論は表現の脆さをなくすのではなく,正確性,信頼性,レイテンシ,コストの新たなトレードオフを明らかにする。
LLMの評価と展開において,表現は第一級インタフェース設計変数として扱われるべきである,と我々は主張する。
関連論文リスト
- Measuring Representation Robustness in Large Language Models for Geometry [7.743292557234699]
幾何学において、同一の問題はユークリッド、座標、ベクトル形式で表すことができる。
既存のベンチマークでは、固定フォーマットの精度が報告されている。
表現対応評価フレームワークGeoRepEvalを提案する。
論文 参考訳(メタデータ) (2026-04-03T11:36:49Z) - Adaptive Problem Generation via Symbolic Representations [16.05958546676182]
本稿では,数学的なタスクにおいて,小さなオープンウェイト言語モデルを改善するために,検証可能な報酬を用いた強化学習のためのトレーニングデータを生成する方法を提案する。
シンボル変数と制約の集合として各問題を表現し、シンボル問題空間で修正を行う。
この表現は、問題構造を正確に制御し、基底構造解の自動生成を可能にし、数学的推論を言語的実現から切り離す。
論文 参考訳(メタデータ) (2026-02-22T13:33:48Z) - From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics [79.81905350372067]
我々は文脈的数学的推論を通してギャップを研究する。
AIMEとMATH-500の問題を2つのコンテキスト設定に再利用するベンチマークであるContextMATHを紹介する。
オープンソースモデルはSGとCSで13、34ポイント減少し、プロプライエタリモデルは13、20ポイント減少している。
論文 参考訳(メタデータ) (2026-01-30T14:56:04Z) - Making Mathematical Reasoning Adaptive [61.45161826629692]
大規模言語モデル(LLM)における適応推論を実現するためのAdaRフレームワークを提案する。
AdaRは可変値によって論理的に等価なクエリを合成し、これらのデータに基づいてRLVRでモデルを訓練し、スプリアス論理をペナルライズする。
実験により, AdaRはロバスト性や一般化を向上し, 数学的推論の大幅な改善を実現していることが示された。
論文 参考訳(メタデータ) (2025-10-06T09:30:05Z) - Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors [11.169118114200307]
大規模言語モデル(LLM)は、強力な数学的問題解決能力を示すが、しばしば訓練分布から構文的に逸脱する問題に失敗する。
モデルがセマンティックに単純だが、不慣れな方法で言い換えられるような問題に対して、慣れ親しんだ推論戦略を誤って適用する、系統的な障害モード、統語的盲点を識別する。
以上の結果から,多くの推論誤差は概念的困難というよりも構造的不整合に起因することが示唆され,構文認識による介入がこれらの帰納的障害を明らかにし緩和する可能性が示唆された。
論文 参考訳(メタデータ) (2025-10-02T09:26:26Z) - A Causal Framework to Quantify the Robustness of Mathematical Reasoning
with Language Models [81.15974174627785]
入力空間における直接的介入に対する頑健さと感度の観点から言語モデルの振舞いについて検討する。
しかし, GPT-3 Davinciモデル(175B)は, 他のGPTモデルと比較して, 頑健さと感度の両面で劇的な改善を実現している。
論文 参考訳(メタデータ) (2022-10-21T15:12:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。