論文の概要: WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
- arxiv url: http://arxiv.org/abs/2607.18084v1
- Date: Mon, 20 Jul 2026 15:52:48 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.692742
- Title: WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
- Title(参考訳): WorldCupArena:フットボール放送における言語モデルとディープリサーチエージェントの微粒化評価
- Authors: Zhaokai Wang, Tianlin Gui, Jiayuan Rao, Shangzhe Di, Yihong Tang, Dingli Liang,
- Abstract要約: We present WorldCupArena, a benchmark for language model and Deep-Research agent。
結果とスコア、おそらくプレイヤーやイベント、統計、そして競技結果を予測する。
予測されたスコアが近いが正確ではない場合に、結果の精度、精度、スコアラインスコアを報告する。
- 参考スコア(独自算出の注目度): 10.941871398755403
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluation, and the same process can be reused for future leagues and cups. Before each match, a model either receives a common evidence package or searches for information itself. It predicts the result and score, likely players and events, match statistics, and the outcome of the competition. After the match, these predictions are compared with the recorded result. We report result accuracy, exact-score accuracy, and a scoreline score that gives some credit when a predicted score is close but not exact, together with scores for the other prediction tasks. Across 104 matches and 13 systems, models with similar result accuracy differ more clearly on detailed predictions. Compared with betting-market and human-fan baselines, the best system shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline. New schedules can be added as they begin, allowing the benchmark to evaluate future models without using outcomes that are already known. Code, prompts, predictions, and evaluation scripts are open sourced at https://github.com/wzk1015/WorldCupArena.
- Abstract(参考訳): キックオフ前にフットボールの試合を予測するには、過去の結果を知ること以上のことが必要である:モデルは変更情報を使用し、回答が利用可能になる前に明確な予測をしなければならない。
We present WorldCupArena, a dynamic benchmark for language model and Deep-Research agent。
2026 FIFAワールドカップは最初の評価であり、このプロセスは将来のリーグやカップで再利用することができる。
各マッチの前に、モデルは共通のエビデンスパッケージを受け取るか、情報自体を検索する。
結果とスコア、おそらくプレイヤーやイベント、統計、そして競技結果を予測する。
試合後、これらの予測は記録された結果と比較される。
予測されたスコアが近いが正確ではない場合、他の予測タスクのスコアとともに、結果の正確さ、正確度、スコアラインスコアを報告する。
104のマッチと13のシステムで、類似した結果の精度を持つモデルは、より詳細な予測でより明確に異なる。
賭け市場やファンベースラインと比較すると、最良のシステムは結果と正確なスコアの精度がわずかに向上するが、スコアラインの利得はより明確である。
新たなスケジュールを追加することで、ベンチマークが既知の結果を用いることなく、将来のモデルを評価することが可能になる。
コード、プロンプト、予測、評価スクリプトはhttps://github.com/wzk1015/WorldCupArenaでオープンソース化されている。
関連論文リスト
- Scaling Open-Ended Reasoning to Predict the Future [56.672065928345525]
我々は、オープンエンドの予測質問の予測を行うために言語モデルを訓練する。
トレーニングデータをスケールアップするために、毎日のニュースで報告されるグローバルイベントから新しい予測質問を合成する。
トレーニングの予測によるキャリブレーションの改善は、一般的なベンチマークで一般化されている。
論文 参考訳(メタデータ) (2025-12-31T18:59:51Z) - Consistency Checks for Language Model Forecasters [54.62507816753479]
予測器の性能を,論理的に異なる質問に対する予測の整合性の観点から測定する。
我々は,一連の基本質問を生成し,これらの質問から整合性チェックをインスタンス化し,予測者の予測を導き,予測の整合性を測定する自動評価システムを構築した。
論文 参考訳(メタデータ) (2024-12-24T16:51:35Z) - Evaluating Soccer Match Prediction Models: A Deep Learning Approach and
Feature Optimization for Gradient-Boosted Trees [0.8009842832476994]
2023年のサッカーの予測チャレンジでは、まず各チームが得点した正確なゴールについて、次に勝利、引き分け、負けの確率について、試合結果の予測が必要とされた。
CatBoost モデルは pi-ratings を特徴として用いており、これは最初は win/draw/loss 確率を計算するのに最適な選択肢として認識されていた。
本研究では,ディープラーニングモデルの性能を評価し,勾配木モデルに最適な特徴セットを決定することを目的とした。
論文 参考訳(メタデータ) (2023-09-26T10:05:46Z) - Predicting Football Match Outcomes with eXplainable Machine Learning and
the Kelly Index [0.0]
フットボールの試合の結果を予測するための機械学習アプローチが開発されている。
このデータセットは、2019-2021シーズンをカバーするプレミアリーグの試合データに由来する。
また、本書の確率をベンチマークすることで、その効果を評価するための投資戦略も考案した。
論文 参考訳(メタデータ) (2022-11-28T19:32:58Z) - Betting the system: Using lineups to predict football scores [0.0]
本稿では,決勝点におけるラインアップの役割を分析し,サッカーにおけるランダム性を低減することを目的とする。
サッカークラブはラインナップに数百万ドルを投資し、個々の統計がより良い結果にどのように変換するかを知ることで投資を最適化することができる。
スポーツの賭けは指数関数的に増加しており、将来を予測することは利益があり、望ましい。
論文 参考訳(メタデータ) (2022-10-12T15:47:42Z) - GCN-WP -- Semi-Supervised Graph Convolutional Networks for Win
Prediction in Esports [84.55775845090542]
本稿では,グラフ畳み込みネットワークに基づくエスポートに対する半教師付き勝利予測モデルを提案する。
GCN-WPはマッチとプレーヤに関する30以上の機能を統合し、近隣のゲームを分類するためにグラフ畳み込みを使用している。
本モデルは,LLの機械学習やスキル評価モデルと比較して,最先端の予測精度を実現する。
論文 参考訳(メタデータ) (2022-07-26T21:38:07Z) - Evaluation of soccer team defense based on prediction models of ball
recovery and being attacked [0.8921166277011345]
本研究では,ボールの回復と攻撃の予測に基づいて,チーム防御を評価する手法を提案する。
45試合のデータを用いて,提案する指標とチームパフォーマンスの関係を検討した。
論文 参考訳(メタデータ) (2021-03-17T13:15:41Z) - Interpretable Real-Time Win Prediction for Honor of Kings, a Popular
Mobile MOBA Esport [51.20042288437171]
本研究では,2段階空間時間ネットワーク(TSSTN)を提案する。
実世界のライブストリーミングシナリオにおける実験結果と応用により,提案したTSSTNモデルは予測精度と解釈可能性の両方において有効であることが示された。
論文 参考訳(メタデータ) (2020-08-14T12:00:58Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。