LLM-SoccerArena Benchmarks AI Sports Predictions
Researchers have introduced LLM-SoccerArena, a prospective live benchmark designed to evaluate how large language models (LLMs) forecast real-world sports events. The platform, available at llm-soccerarena.com, provides a public open-source system that automatically records timestamped, schema-validated forecasts of unresolved events, along with prompts, model versions, tool traces, and costs. The benchmark uses a factorial design with tournament-related questions, such as which team will win, to test LLMs' ability to synthesize information under uncertainty. This addresses the limitation of existing static and retrospective benchmarks, which cannot assess predictive capabilities for future events. The project aims to advance understanding of LLMs' decision-making in uncertain, real-world contexts.
Key facts
- LLM-SoccerArena is a prospective live benchmark for LLMs.
- It evaluates LLMs on forecasting real-world sports events.
- The platform is open-source and available at llm-soccerarena.com.
- It records timestamped, schema-validated forecasts with metadata.
- The benchmark uses a factorial design with tournament-related questions.
- Existing benchmarks are static and retrospective, limiting evaluation of predictive capabilities.
- The project addresses how LLMs synthesize information to predict future events.
- The benchmark includes prompts, model versions, tool traces, and costs.
Entities
Institutions
- arXiv