Fugu-MT 論文翻訳(概要): TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving

論文の概要: TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving

arxiv url: http://arxiv.org/abs/2504.15780v2
Date: Fri, 29 Aug 2025 08:38:35 GMT
ステータス: 翻訳完了
システム内更新日: 2025-09-01 17:44:08.732637
Title: TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
Title（参考訳）: TrustGeoGen: 信頼できるマルチモーダル幾何学的問題解決のための形式検証データエンジン
Authors: Daocheng Fu, Jianlong Chen, Renqiu Xia, Zijun Chen, Qi Liu, Yuan Feng, Hongbin Zhou, Renrui Zhang, Shiyang Feng, Peng Gao, Hongyuan Zha, Junchi Yan, Botian Shi, Yu Qiao, Bo Zhang,
Abstract要約: TrustGeoGenは、標準的で信頼性の高いベンチマークを確立するために、正式に検証された幾何問題を生成するデータエンジンである。 1)ダイアグラム,テキスト,ステップバイステップのソリューションの生成を同期するマルチモーダルアライメント,2)すべての推論パスがルール準拠であることを保証する形式検証,3)接続思考,ブリッジング,ヒューマンライクな論理ステップとの論理的推論,4)複数のソリューションと自己回帰バックトラックを備えた多種多様な問題を生成できるTextitGeoExploreシリーズアルゴリズム。
参考スコア（独自算出の注目度）: 106.04001249574786
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Mathematical geometric problem solving (GPS) demands verifiable logical coherence and multimodal reasoning capabilities. While large language models (LLMs) have shown rapid progress in GPS, their advancement is hindered by the lack of reliable benchmarks and systematic methodologies. A critical challenge is the inherent hallucination in LLMs, which leads to synthetic GPS datasets that are often noisy, unverified, and self-contradictory. To address this, we introduce TrustGeoGen, a data engine that generates formally verified geometric problems to establish a principled and trustworthy benchmark. Our engine integrates four key innovations: 1) Multimodal Alignment, which synchronizes the generation of diagrams, text, and step-by-step solutions; 2) Formal Verification, ensuring all reasoning paths are rule-compliant; 3) Connection Thinking, bridging formal deduction with human-like logical steps; and 4) our \textit{GeoExplore} series algorithms, which produce diverse problem variants with multiple solutions and self-reflective backtracking. Using this engine, we create the GeoTrust-200K dataset and the corresponding GeoTrust-test benchmark, both with guaranteed cross-modal integrity. Experiments reveal that state-of-the-art models achieve only 45.83\% accuracy on GeoTrust-test, highlighting its significant challenge. Furthermore, training on our synthesized data substantially improves model performance on GPS tasks, with strong generalization to out-of-domain (OOD) benchmarks. Our code and data are available at https://github.com/Alpha-Innovator/TrustGeoGen
Abstract（参考訳）: 数学的幾何学的問題解決(GPS)は、検証可能な論理コヒーレンスとマルチモーダル推論能力を必要とする。大規模言語モデル(LLM)はGPSの急速な進歩を示す一方で、信頼性の高いベンチマークや体系的な手法の欠如によってその進歩は妨げられている。重要な課題は、LLMの固有の幻覚であり、しばしばうるさい、証明されていない、自己矛盾的な合成GPSデータセットに繋がる。この問題を解決するために、TrustGeoGenというデータエンジンを導入します。私たちのエンジンは4つの重要なイノベーションを統合しています。 1) 図形,テキスト及びステップバイステップのソリューションの生成を同期させるマルチモーダルアライメント 2) 形式的検証,すべての理由付けパスが規則に準拠していることを保証する。 3)接続思考,人間的な論理的ステップによる形式的推論,及び 4) 複数解と自己回帰バックトラックを用いた多種多様な問題変種を生成する。このエンジンを用いて,GeoTrust-200Kデータセットと対応するGeoTrust-testベンチマークを作成する。実験の結果、現状のモデルはGeoTrust-testで45.83倍の精度しか達成していないことが判明した。さらに, 合成データのトレーニングにより, GPSタスクのモデル性能が大幅に向上し, ドメイン外ベンチマーク(OOD)への強力な一般化が期待できる。私たちのコードとデータはhttps://github.com/Alpha-Innovator/TrustGeoGenで公開されています。

論文の概要: TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving

関連論文リスト