論文の概要: Continual Graph Memory for Mathematical Research Agents
- arxiv url: http://arxiv.org/abs/2610.02945v1
- Date: Fri, 02 Oct 2026 07:38:09 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-06 00:14:30.255857
- Title: Continual Graph Memory for Mathematical Research Agents
- Title(参考訳): 数学研究エージェントのための連続グラフメモリ
- Abstract要約: Ansatzはグラフベースの、進化可能な、クロスプロブレムな数学的研究メモリシステムである。
証明探索プロセス全体を整理し、過去の問題の探索軌跡から情報を再利用する。
アンサッツは10の研究課題の全てを閉鎖し、長期の数学的探索を継続し再開する能力を示した。
- 参考スコア(独自算出の注目度): 95.750048838182
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Using frontier agent harnesses to tackle mathematical research problems has emerged as an effective means of advancing mathematics. However, solving frontier problems in mathematics may require a massive number of agents working in parallel for extended periods to construct proofs, thereby generating an enormous volume of intermediate proof results. Organizing these intermediate results throughout a long-horizon proof-search process and reusing knowledge gained from prior explorations remain major challenges. We present Ansatz, a mathematical research agent built around Continual Graph Memory, a graph-based, evolvable, cross-problem mathematical research memory system that explicitly organizes the entire proof search process and reuses information from exploration trajectories of previous problems. Specifically, we develop a unified graph memory that represents all intermediate exploration results, including facts, plans, and counterexamples, together with edges that explicitly represent the relationships among them; dependency-aware retrieval supplies precisely targeted local context; an evidence-sensitive curator updates the research frontier and distills lessons from prior attempts; and scoped recall surfaces earlier statements and negative findings for local re-proving rather than uncritical reuse. Experiments cover runs across all ten First Proof Second Batch problems, together with four component studies. Ansatz reports closure on all ten research tasks, demonstrating its ability to sustain and resume long-horizon mathematical search. Beyond these problems, Ansatz also produces solutions to the Jamison caterpillar conjecture and Erdős Problems 289, 348, and 488 without human intervention, and makes partial progress on several open problems, illustrating its strong ability to solve open mathematical research problems.
- Abstract(参考訳): 数学研究問題に対処するためのフロンティアエージェントハーネスは、数学を前進させる効果的な手段として登場した。
しかし、数学におけるフロンティア問題の解法は、証明を構築するために長い期間並列に働く大量のエージェントを必要とし、結果として大量の中間証明結果を生成する。
これらの中間結果を長期にわたる実証研究プロセスを通じて整理し、先行調査から得られた知識を再利用することは大きな課題である。
グラフベースで進化可能でクロスプロブレムな数学的研究用メモリシステムであるContinual Graph Memoryを中心に構築された数学研究エージェントAnsatzについて述べる。
具体的には、事実、計画、反例を含む全ての中間探索結果を表す統一的なグラフメモリと、それら間の関係を明示するエッジと、それら間の関係を的確に表わすエッジ、エビデンスに敏感なキュレーターが研究フロンティアを更新し、以前の試みから教訓を抽出する、スコープ化されたリコールサーフェスと、非クリティカルな再利用ではなく、局所的な再実装のための負の発見を含む、統一的なグラフメモリを開発する。
実験は、10の第一証明第二バッチ問題と4つのコンポーネント研究をカバーしている。
アンサッツは10の研究課題の全てを閉鎖し、長期の数学的探索を継続し再開する能力を示した。
これらの問題以外にも、アンザッツはJamison caterpillar conjecture(英語版)やエルデシュ問題(英語版) 289, 348, 488(英語版)の解を人間の介入なしに生成し、オープンな数学的研究問題を解く強力な能力を示している。
関連論文リスト
- The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models [65.00652251189636]
大規模言語モデル(LLM)における数学的理解を体系的に研究する第一歩を踏み出す。
本研究では,発見,生成,消化,実行の4つの異なる側面に沿って数学的推論を評価する新しいベンチマークであるhleiを提案する。
次に、学生モデルにプリミティブ誘導推論を選択的に転送するプリミティブプリミティブ自己蒸留フレームワークであるabsを紹介する。
論文 参考訳(メタデータ) (2026-10-01T17:59:32Z) - AI and Human Approaches to Mathematical Problem Solving [2.799791426680068]
この研究は、公的なAI研究のアカウントと、そのような11の問題についての人間の文献を比較した。
人間のコーパスには58の論文が含まれており、後にAIソースが解決、不承認、あるいは実質的な進歩として報告したのと同じ数学的標的に対処している。
AIアカウントは問題を閉じて再結合することに重点を置いているのに対し、数学論文では、結果が累積的な知識になるための手順、限界、研究の機会をより広範囲に文書化している。
論文 参考訳(メタデータ) (2026-09-15T19:40:34Z) - From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier [109.93387172162984]
AI4Mathシステムの次の飛躍は、事前に定義された問題解決者から研究エージェントへの決定的なシフトを必要とする。
この分野の体系的なレビューを行い、データセット、自動形式化、証明合成について紹介する。
論文 参考訳(メタデータ) (2026-07-08T17:46:36Z) - Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory [7.10978537067254]
Danusは、共有事実グラフを中心とした研究レベルの数学的推論のためのオーケストレーションシステムである。
検証された各事実は証明と論理的依存関係と共に格納され、システムは長い引数を構築できる。
論文 参考訳(メタデータ) (2026-07-07T16:11:30Z) - CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions [7.449578020792231]
我々は、MIT PRIMES--Art of Problem Solving (AoPS)プログラムから164のエキスパートアノテートプログレスチェーンのデータセットであるCrowdMathを紹介する。
各チェーンは、オープンプロブレムステートメントから完成した証明まで、多人数のフォーラムディスカッションをトレースする。
モデルは次のポスト予測において83~88%の精度を達成し、数学的議論の局所的な流れに従うことができることを示唆している。
論文 参考訳(メタデータ) (2026-06-02T20:38:39Z) - HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification [54.06301039725887]
計算および応用数学において8つの領域にまたがる100以上の未解決問題のベンチマークであるHorizonMathを紹介する。
我々のベンチマークは、発見が困難であり、意味のある数学的洞察を必要とする問題のクラスをターゲットにしているが、検証は計算的に効率的で簡単なものである。
論文 参考訳(メタデータ) (2026-03-16T17:59:53Z) - UniGeo: Unifying Geometry Logical Reasoning via Reformulating
Mathematical Expression [127.68780714438103]
計算と証明の2つの主要な幾何学問題は、通常2つの特定のタスクとして扱われる。
我々は4,998の計算問題と9,543の証明問題を含むUniGeoという大規模統一幾何問題ベンチマークを構築した。
また,複数タスクの幾何変換フレームワークであるGeoformerを提案し,計算と証明を同時に行う。
論文 参考訳(メタデータ) (2022-12-06T04:37:51Z) - GeoQA: A Geometric Question Answering Benchmark Towards Multimodal
Numerical Reasoning [172.36214872466707]
我々は、テキスト記述、視覚図、定理知識の包括的理解を必要とする幾何学的問題を解くことに注力する。
そこで本研究では,5,010の幾何学的問題を含む幾何学的質問応答データセットGeoQAを提案する。
論文 参考訳(メタデータ) (2021-05-30T12:34:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。