論文の概要: Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation
- arxiv url: http://arxiv.org/abs/2606.10833v1
- Date: Tue, 09 Jun 2026 13:20:15 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-10 15:40:58.516836
- Title: Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation
- Title(参考訳): VLMはエンジニアに似ていますか?ベンチマークと段階評価
- Authors: Syed Wasiq, Syed Mohamad Tawseeq, Yashwant Pravinrao Bangde, Debaditya Roy,
- Abstract要約: EngVQAは、696の問題を含む5つのエンジニアリング分野にわたるエンジニアリング推論を評価するためのマルチモーダルベンチマークである。
VLM生成ソリューション評価のための8段階自動評価フレームワークを提案する。
本研究は,マルチモーダルエンジニアリング推論システムの信頼性評価におけるプロセス指向評価の重要性を強調した。
- 参考スコア(独自算出の注目度): 3.2128810211809196
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reasoning remains largely unexplored. Unlike general visual question answering, engineering problem solving requires interpreting technical diagrams, selecting governing physical principles, and maintaining physically consistent multi-step reasoning. These capabilities are increasingly important for AI systems used in engineering education, scientific assistance, and technical decision-making, where reasoning failures may produce physically invalid yet superficially plausible solutions. Existing benchmarks primarily evaluate final answers and provide limited assessment of intermediate reasoning processes. We introduce EngVQA, a multimodal benchmark for evaluating engineering reasoning across 5 engineering subjects containing 696 problems. We introduce an 8-stage automatic evaluation framework for assessing VLM-generated solutions. The framework independently evaluates each stage of the solution, enabling fine-grained analysis of reasoning failures. We benchmark multiple state-of-the-art open and closed source VLMs on our evaluation framework and demonstrate substantial limitations in current engineering reasoning capabilities. Human evaluation shows strong agreement with our automated framework, achieving a Pearson correlation of 0.975 and a mean absolute error of 0.67 on a 10-point grading scale. Our results highlight the importance of process-oriented evaluation for reliable assessment of multimodal engineering reasoning systems.
- Abstract(参考訳): VLM(Vision-Language Models)は、一般的なマルチモーダル推論ベンチマークにおいて強力な性能を示すが、工学的推論を行う能力はいまだ明らかにされていない。
一般的な視覚的質問応答とは異なり、工学的な問題解決には、技術的な図の解釈、物理的な原則の選択、物理的に一貫した多段階推論の維持が必要である。
これらの能力は、工学教育、科学支援、技術決定に使用されるAIシステムにとってますます重要になっている。
既存のベンチマークは主に最終回答を評価し、中間推論プロセスの限定的な評価を提供する。
本稿では,696問題を含む5つの工学科目を対象に,工学的推論を評価するためのマルチモーダルベンチマークであるEngVQAを紹介する。
VLM生成ソリューション評価のための8段階自動評価フレームワークを提案する。
このフレームワークはソリューションの各段階を独立に評価し、推論失敗のきめ細かい分析を可能にする。
我々は、評価フレームワーク上で、最先端のオープンソースVLMとクローズドソースVLMをベンチマークし、現在のエンジニアリング推論能力の大幅な制限を実証する。
Pearsonの相関は0.975であり,平均絶対誤差は0.67である。
本研究は,マルチモーダルエンジニアリング推論システムの信頼性評価におけるプロセス指向評価の重要性を強調した。
関連論文リスト
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning [57.868248683256574]
PRISM-Physicsはプロセスレベルの評価フレームワークであり、複雑な物理推論問題のベンチマークである。
解は公式の有向非巡回グラフ(DAG)として表される。
その結果,評価フレームワークは人的専門家のスコアと一致していることがわかった。
論文 参考訳(メタデータ) (2025-10-03T17:09:03Z) - MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs [55.20845457594977]
大規模言語モデル(LLM)は、問題解決と意思決定の能力の向上を示している。
本稿ではメタ推論技術を必要とするプロセスベースのベンチマークMR-Benを提案する。
メタ推論のパラダイムは,システム2のスロー思考に特に適しています。
論文 参考訳(メタデータ) (2024-06-20T03:50:23Z) - Evaluating Mathematical Reasoning Beyond Accuracy [50.09931172314218]
推論ステップの品質を評価するための新しい方法論であるReasonEvalを紹介します。
ReasonEvalはメタ評価データセットのベースライン手法よりも一貫して優れていることを示す。
我々は、ReasonEvalがデータ選択において重要な役割を果たすことを観察する。
論文 参考訳(メタデータ) (2024-04-08T17:18:04Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。