論文の概要: Objects Without Morphisms: What LLMs for Mathematics Do Not Represent
- arxiv url: http://arxiv.org/abs/2610.03551v1
- Date: Fri, 02 Oct 2026 16:33:32 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-06 00:14:30.478786
- Title: Objects Without Morphisms: What LLMs for Mathematics Do Not Represent
- Title(参考訳): モルヒズムのない物体:数学のLLMが表現しないもの
- Abstract要約: 大規模言語モデル (LLM) は、競争数学のエキスパートレベルの性能に到達した。
周辺地域の方言間での文の翻訳という,外部基準が存在しない問題について検討する。
我々は、個別のブラインドキューで真実、内容、スコープを符号化する機器を導入する。
- 参考スコア(独自算出の注目度): 27.64177147171874
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Large language models (LLMs) have reached expert-level performance on competition mathematics largely through the volume of search placed around them: candidate solutions are sampled in quantity and retained only when an external criterion accepts them. Such a procedure improves the outcome that survives it while leaving untouched what the model represents. We examine that question where no external criterion exists: translating statements between the dialects of neighbouring subfields, where fidelity turns on the level of generality at which content is asserted. The source leaves that level implicit in its vocabulary, so a faithful translation must recover it from the relation between the theories. We introduce an instrument that codes truth, content and scope in separate blind queues, with a judge-free measure of whether a rewrite states the hypothesis implicit in its source, and establish its sensitivity with a planted-positive control. Across seven models from four families, translating towards the general framing widens the domain of quantification in 60.6% of rewrites and narrows it in none; translating towards the concrete framing narrows it in 28.3% and widens it in 0.3%. The hypothesis that would prevent it is stated in 21.6% of model rewrites and 4.2% of human statements. Capability does not govern the asymmetry: it appears in every model tested, and the most capable widens least. It replicates on the half of the benchmark held out by a pre-registered rule, and on statements written by mathematicians. Instructing a model to state every hypothesis it requires raises that rate but not its sensitivity to direction. We argue that these systems have acquired an object-level correspondence between subfield vocabularies without the constraint under which a translation between theories carries hypotheses to hypotheses.
- Abstract(参考訳): 大規模言語モデル (LLMs) は、主にその周辺に配置された探索量を通じて、競争数学のエキスパートレベルのパフォーマンスに達している。
このようなプロシージャは、モデルが表現しているものに触れることなく、それを生き残る結果を改善する。
我々は、外部の基準が存在しない問題について考察する: 近隣の亜分野の方言間で文を翻訳し、その内容が主張される一般性のレベルに忠実さが移る。
情報源は語彙にそのレベルを暗黙に残しているため、忠実な翻訳は理論の関係からそれを取り戻す必要がある。
そこで我々は,真理,内容,範囲を個別のブラインドキューで符号化し,書き直しが仮説を暗黙に表現するかどうかを判断自由度で判断し,その感度を植木正の制御で確立する手法を提案する。
4つの家系の7つのモデルにまたがって、一般的なフレーミングに向けての翻訳は、60.6%の書き直しで定量化の領域を広げ、それを全く狭め、コンクリートフレーミングへの翻訳は28.3%の絞り込み、0.3%の幅で拡大する。
これを防ぐための仮説は21.6%のモデル書き換えと4.2%のヒューマンステートメントで述べられている。
能力は非対称性を支配せず、テストされた全てのモデルに現れ、最も有能なものは最小限に拡大する。
これは、事前登録されたルールによって保持されるベンチマークの半分と、数学者によって書かれたステートメントを再現する。
モデルに要求される全ての仮説を述べるように指示すると、その速度は上昇するが、方向に対する感度は上昇しない。
これらの系は、理論間の翻訳が仮説を導く制約を伴わずに、サブフィールド語彙間のオブジェクトレベル対応を得たと論じる。
関連論文リスト
- Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models [0.0]
我々はこの疑問を、非アービタージュのレンズを通して「理解」を定義することで測定可能とする。
計算的に有界なトレーダーが保証された利益を抽出できない場合、モデルは語彙をある程度理解する。
Arbitrは、敵対的トレーダーが論理的不整合のモデルをペナルティ化するトレーニングフレームワークである。
論文 参考訳(メタデータ) (2026-09-30T09:07:44Z) - Not What You Meant: Can LLMs Follow a Specified Negation Semantics? [4.575311337142799]
否定はドメイン間の一様解釈を持たない。
法的、規制的、医学的理由づけにおいて、意図された解釈は、効力のある読みに依存する。
我々は,否定的大言語モデルの読み出しをデフォルトで採用する方法について検討する。
論文 参考訳(メタデータ) (2026-09-23T08:14:47Z) - Message capacity and claim wording set the transition points of collective truth-finding in language-model networks [0.0]
集合の運命は、クレームのワードセットのしきい値と、遷移点を設定するメッセージキャパシティの2つの単一エージェント測定によって決定される。
論文 参考訳(メタデータ) (2026-09-15T14:46:31Z) - Transformed in Translation: Two-Stage Structural Uncertainty in LLM-Based Scientific Autoformalization [0.0]
オートフォーマル化はアカウントを実行可能な数学に変えるが、実行可能コードはどのモデルが構築されたかは決着しない。
応答法則を生成する形式化器と,その法則を軌道に変換する再帰性という,構造的不確実性の2つの源について検討する。
論文 参考訳(メタデータ) (2026-09-13T21:47:46Z) - A Formal Limitation on Learning Human Language From Textual Corpora [54.22186267212778]
我々は,デコーダが発話の表現から話者の意図した意味を回復する確率に基づいて,上位境界を導出する。
人工言語、マンダリンゼロ代名詞分解、色基準の実験は、この理論を支持する実証的な証拠を提供する。
論文 参考訳(メタデータ) (2026-08-28T17:38:24Z) - Solving Inequality Proofs with Large Language Models [42.667163027148916]
不等式証明は様々な科学・数学分野において不可欠である。
これにより、大きな言語モデル(LLM)の需要が高まるフロンティアとなる。
我々は、Olympiadレベルの不平等を専門家が計算したデータセットであるIneqMathをリリースした。
論文 参考訳(メタデータ) (2025-06-09T16:43:38Z) - Uncertainty in Language Models: Assessment through Rank-Calibration [65.10149293133846]
言語モデル(LM)は、自然言語生成において有望な性能を示している。
与えられた入力に応答する際の不確実性を正確に定量化することは重要である。
我々は、LMの確実性と信頼性を評価するために、Rank$-$Calibration$と呼ばれる斬新で実用的なフレームワークを開発する。
論文 参考訳(メタデータ) (2024-04-04T02:31:05Z) - Paloma: A Benchmark for Evaluating Language Model Fit [112.481957296585]
言語モデル (LM) の評価では、トレーニングから切り離されたモノリシックなデータに難易度が報告されるのが一般的である。
Paloma(Perplexity Analysis for Language Model Assessment)は、546の英語およびコードドメインに適合するLMを測定するベンチマークである。
論文 参考訳(メタデータ) (2023-12-16T19:12:45Z) - On the Usefulness of Embeddings, Clusters and Strings for Text Generator
Evaluation [86.19634542434711]
Mauveは、弦上の2つの確率分布間の情報理論のばらつきを測定する。
我々は,Mauveが誤った理由で正しいことを示し,新たに提案された分岐はハイパフォーマンスには必要ないことを示した。
テキストの構文的およびコヒーレンスレベルの特徴を符号化することで、表面的な特徴を無視しながら、文字列分布に対するクラスタベースの代替品は、単に最先端の言語ジェネレータを評価するのに良いかもしれない、と結論付けています。
論文 参考訳(メタデータ) (2022-05-31T17:58:49Z) - A Weaker Faithfulness Assumption based on Triple Interactions [89.59955143854556]
より弱い仮定として, 2$-adjacency faithfulness を提案します。
より弱い仮定の下で適用可能な因果発見のための音方向規則を提案する。
論文 参考訳(メタデータ) (2020-10-27T13:04:08Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。