論文の概要: BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation
- arxiv url: http://arxiv.org/abs/2605.28994v1
- Date: Wed, 27 May 2026 18:51:00 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-30 02:45:55.23914
- Title: BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation
- Title(参考訳): BEAMS: モデリングとシミュレーションのためのAIのベンチマークと評価
- Authors: Sara Metcalf, William Schoenberg,
- Abstract要約: BEAMSイニシアティブは、モデリングとシミュレーションのためのAIツールの開発を、責任と倫理的な形式へと導くことを目的としている。
このイニシアチブは、オープンなデジタルおよび組織インフラを使用して、モデリングとシミュレーションのためのAIツールを協調的に評価する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable. Tools that can automate aspects of modeling practice must complement human expertise, not replace it. The BEAMS Initiative aims to guide the development of AI tools for modeling and simulation toward forms that are responsible and ethical by establishing benchmarks for human centered modeling and simulation practices. The initiative uses open digital and organizational infrastructure to collaboratively evaluate AI tools for modeling and simulation. The open source sd ai project hosted by the initiative establishes transparency and enables contributions to be shared broadly. A steering group focuses on prioritizing potential benchmarks, while a technical group focuses on implementing the benchmarks in the form of automated tests. Tests for several distinct categories of evaluation have been implemented and applied to AI tools that support qualitative model building, quantitative model building, and model discussion. These include tests for causal translation, model iteration, causal reasoning, conformance, model behavior explanation, suggested model building steps, and suggested model fixes. When engines from the sd ai project are coupled with different LLMs, their performance on these evaluations reveals variability across different AI tools. The evaluations implemented by the initiative demonstrate that AI enabled modeling tools perform better at discussion and basic qualitative tasks than with causal reasoning and quantitative error fixing. No single LLM dominates across engine types, highlighting the importance of specific tasks and tradeoffs between speed and accuracy. Ongoing efforts of the initiative aim to incorporate benchmarks that address concerns about bias by considering alternative perspectives and human centered use cases.
- Abstract(参考訳): 現実世界の意思決定をサポートするAIツールは、レコメンデーションを通知し、解釈可能なシミュレーションモデルを構築することができる必要がある。
モデリングプラクティスの側面を自動化できるツールは、それを置き換えるのではなく、人間の専門知識を補完しなければならない。
BEAMSイニシアティブは、人間中心のモデリングとシミュレーションのプラクティスのベンチマークを確立することで、責任と倫理性を持つ形式に向けて、モデリングとシミュレーションのためのAIツールの開発をガイドすることを目的としている。
このイニシアチブは、オープンなデジタルおよび組織インフラを使用して、モデリングとシミュレーションのためのAIツールを協調的に評価する。
このイニシアチブが主催するオープンソースのsd aiプロジェクトは透明性を確立し、コントリビューションを広く共有することを可能にする。
ステアリンググループは潜在的なベンチマークの優先順位付けに重点を置いており、テクニカルグループは自動テストの形でベンチマークを実装することに重点を置いている。
いくつかの異なる評価カテゴリのテストが実装され、定性的モデル構築、定量的モデル構築、モデル議論をサポートするAIツールに適用されている。
これには、因果翻訳のテスト、モデルイテレーション、因果推論、適合性、モデル振る舞いの説明、モデル構築ステップの提案、モデル修正の提案が含まれる。
sd aiプロジェクトのエンジンが異なるLLMと結合されると、これらの評価のパフォーマンスは、異なるAIツール間でのばらつきを明らかにします。
このイニシアチブが実施した評価によると、AIを有効にしたモデリングツールは、因果推論や量的誤り修正よりも、議論や質的なタスクにおいて優れたパフォーマンスを発揮する。
LLMはエンジンタイプで支配的な存在ではなく、特定のタスクの重要性と速度と精度のトレードオフを強調している。
このイニシアティブの取り組みは、代替的な視点と人間中心のユースケースを考慮してバイアスに関する懸念に対処するベンチマークを統合することを目的としている。
関連論文リスト
- Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content [71.46991494014382]
本稿では,AI生成画像におけるリアリズムと妥当性の詳細な評価のための新しいデータセットであるQ-Realを紹介する。
Q-Realは、人気のあるテキスト・ツー・イメージ・モデルによって生成される3,088のイメージで構成されている。
そこで本研究では,Q-Real Benchを2つの課題,すなわち判断と推論による根拠付けに基づいて評価する。
論文 参考訳(メタデータ) (2025-11-21T02:43:17Z) - Modèles de Substitution pour les Modèles à base d'Agents : Enjeux, Méthodes et Applications [0.0]
エージェントベースモデル(ABM)は、局所的な相互作用から生じる創発的な現象を研究するために広く用いられている。
ABMの複雑さは、リアルタイム意思決定と大規模シナリオ分析の可能性を制限する。
これらの制限に対処するため、サロゲートモデルはスパースシミュレーションデータから近似を学習することで効率的な代替手段を提供する。
論文 参考訳(メタデータ) (2025-05-17T08:55:33Z) - ToolACE-DEV: Self-Improving Tool Learning via Decomposition and EVolution [77.86222359025011]
ツール学習のための自己改善フレームワークであるToolACE-DEVを提案する。
まず、ツール学習の目的を、基本的なツール作成とツール利用能力を高めるサブタスクに分解する。
次に、軽量モデルによる自己改善を可能にする自己進化パラダイムを導入し、高度なLCMへの依存を減らす。
論文 参考訳(メタデータ) (2025-05-12T12:48:30Z) - Benchmarks as Microscopes: A Call for Model Metrology [76.64402390208576]
現代の言語モデル(LM)は、能力評価において新たな課題を提起する。
メトリクスに自信を持つためには、モデルミアロジの新たな規律が必要です。
論文 参考訳(メタデータ) (2024-07-22T17:52:12Z) - QualEval: Qualitative Evaluation for Model Improvement [82.73561470966658]
モデル改善のための手段として,自動定性評価による定量的スカラー指標を付加するQualEvalを提案する。
QualEvalは強力なLCM推論器と新しいフレキシブルリニアプログラミングソルバを使用して、人間の読みやすい洞察を生成する。
例えば、その洞察を活用することで、Llama 2モデルの絶対性能が最大15%向上することを示す。
論文 参考訳(メタデータ) (2023-11-06T00:21:44Z) - A Model-Driven Engineering Approach to Machine Learning and Software
Modeling [0.5156484100374059]
モデルは、ソフトウェア工学(SE)と人工知能(AI)のコミュニティで使われている。
主な焦点はIoT(Internet of Things)とCPS(Smart Cyber-Physical Systems)のユースケースである。
論文 参考訳(メタデータ) (2021-07-06T15:50:50Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。