論文の概要: LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
- arxiv url: http://arxiv.org/abs/2607.24573v1
- Date: Mon, 27 Jul 2026 15:38:29 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-28 22:34:15.483289
- Title: LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
- Title(参考訳): LLM-SoccerArena:スポーツにおける実世界の予測に関するLLMのベンチマーク
- Authors: Jonas Schröder, Jonas Schweisthal, Oliver Müller, Markus Weinmann, Stefan Feuerriegel,
- Abstract要約: 大規模言語モデル(LLM)は、不確実な将来の出来事に関する決定をますます支持しているが、実際の結果を予測する能力の評価は難しいままである。
LLM-SoccerArenaは,LLMが実際のスポーツイベントを予測し,その結果が分かる前に評価するライブベンチマークである。
- 参考スコア(独自算出の注目度): 29.80986348128258
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLMs generated forecasts for all 104 matches and 15 tournament-related questions. We provide a detailed analysis of model performance across information access, prompting strategy, and forecast horizon. As a result, LLM-SoccerArena provides new evidence about the forecasting performance of state-of-the-art LLMs. For example, LLMs with web access outperform those without, but only by a small margin (i.e., a 0.023 improvement in Brier score). Overall, LLM-SoccerArena provides a flexible, open-source platform for prospective benchmarking of unresolved events. LLM-SoccerArena will be continuously updated, and can be directly applied to future national and international tournaments and league competitions.
- Abstract(参考訳): 大規模言語モデル(LLM)は、不確実な将来の出来事に関する決定をますます支持しているが、実際の結果を予測する能力の評価は難しいままである。
特に、既存のベンチマークは典型的には静的かつ振り返りであり、したがってLLMによってどのように情報が合成され、不確実な将来の事象を予測するかをテストすることはできない。
LLM-SoccerArena (https://llm-soccerarena.com) は,LLMが実際のスポーツイベントを予測して結果がわかる前にどの程度の精度で予測できるかを評価する,先進的なライブベンチマークである。
LLM-SoccerArenaは、(1)ライブベンチマークプロトコル、(2)オープンソースプラットフォーム、(3)トーナメント関連の質問(例えば、どのチームが勝つか)とともに、因子的ベンチマーク設計を提供する。
LLM-SoccerArenaは、プロンプト、モデルバージョン、ツールトレース、コストとともに、未解決イベントのタイムスタンプ、スキーマ検証された予測を自動的に記録する。
因子設計は、(1)モデルバージョン(例: GPT-5.5、Claude Opus 4.8)、(2)情報アクセス、(3)プロンプト戦略、(4)予測水平線である。
LLM-SoccerArenaは,FIFAワールドカップ2026の大規模評価を通じて,104試合中7試合,トーナメント関連15問の予測を作成した。
本稿では,情報アクセス,戦略の推進,予測地平線におけるモデル性能の詳細な解析を行う。
その結果、LLM-SoccerArenaは最先端のLLMの予測性能に関する新たな証拠を提供する。
例えば、ウェブアクセスを持つLLMは、それ以外のものよりも優れていますが、小さなマージン(ブリアスコアの0.023の改善)によってのみパフォーマンスが向上します。
LLM-SoccerArenaは、未解決イベントの予測ベンチマークのための柔軟なオープンソースプラットフォームを提供する。
LLM-SoccerArenaは継続的に更新され、将来の全国および国際大会やリーグ大会に直接適用される。
関連論文リスト
- Rethinking the Role of LLMs in Time Series Forecasting [15.951870420397682]
大規模言語モデル (LLM) は時系列予測 (TSF) に導入され、数値信号以外の文脈知識が組み込まれている。
このような結論は,限られた評価設定に起因し,大規模に保たないことを示す。
以上の結果から,emphLLM4TSでは予測性能が向上し,ドメイン間の一般化が著しく向上することが示唆された。
論文 参考訳(メタデータ) (2026-02-16T13:39:09Z) - LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena [25.304644327116975]
大規模言語モデル(LLM)は、将来の事象を予測するために、インターネットスケールのデータに基づいて訓練されている。
本稿では,LLMの予測知能について系統的に検討する。
LLM-as-a-Prophetによる優れた予測知能の実現に向けた重要なボトルネックを明らかにする。
論文 参考訳(メタデータ) (2025-10-20T15:20:05Z) - An Empirical Study of Many-to-Many Summarization with Large Language Models [82.10000188179168]
大規模言語モデル(LLM)は強い多言語能力を示しており、実アプリケーションでM2MS(Multi-to-Many summarization)を実行する可能性を秘めている。
本研究は,LLMのM2MS能力に関する系統的研究である。
論文 参考訳(メタデータ) (2025-05-19T11:18:54Z) - Are Large Language Models Strategic Decision Makers? A Study of Performance and Bias in Two-Player Non-Zero-Sum Games [56.70628673595041]
大規模言語モデル (LLM) は現実世界での利用が増えているが、その戦略的意思決定能力はほとんど探索されていない。
本研究は,Stag Hunt と Prisoner Dilemma のカノニカルゲーム理論2人プレイヤ非ゼロサムゲームにおける LLM の性能とメリットについて検討する。
GPT-3.5, GPT-4-Turbo, GPT-4o, Llama-3-8Bの構造化評価は, これらのゲームにおいて決定を行う場合, 以下の系統的バイアスの少なくとも1つの影響を受けていることを示す。
論文 参考訳(メタデータ) (2024-07-05T12:30:02Z) - Large Language Models: A Survey [66.39828929831017]
大規模言語モデル(LLM)は、広範囲の自然言語タスクにおける強力なパフォーマンスのために、多くの注目を集めている。
LLMの汎用言語理解と生成能力は、膨大なテキストデータに基づいて数十億のモデルのパラメータを訓練することで得られる。
論文 参考訳(メタデータ) (2024-02-09T05:37:09Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。