ARTFEED — Contemporary Art Intelligence

SportD: Evaluating Vision-Language Models' Strategic Decision-Making in Soccer

ai-technology · 2026-08-13

A recent study presents SportD, a dataset and evaluation framework designed to assess vision-language models (VLMs) in dynamic settings, specifically utilizing professional soccer as its testing ground. Available on arXiv (ID 2607.14616), the dataset comprises 1,415 decision-making scenarios from both men's and women's soccer games. The research reveals that current VLMs achieve optimal action selection merely 30% of the time, which is inferior to human capabilities. These models tend to opt for safer actions that yield lower variance and value, hindering goal advancement. Although VLMs are skilled at predicting action success, they mistakenly equate likelihood with value, preferring probable actions over those with greater strategic benefit. This study underscores the shortcomings of VLMs in evaluating value and risk, establishing a benchmark for future enhancements.

Key facts

  • SportD is a dataset and evaluation for testing VLMs' strategic decision-making in soccer.
  • The dataset includes 1,415 decision scenarios from professional men's and women's soccer games.
  • VLMs choose the optimal action only around 30% of the time, less than humans.
  • Models prefer safer, lower-variance, lower-value actions that make less progress toward goal.
  • Frontier VLMs place the highest-success-probability action among top choices in 83-92% of cases.
  • VLMs systematically conflate likelihood with value, assigning higher value to more likely actions.
  • The paper is available on arXiv with ID 2607.14616.
  • The study uses soccer as an objective testbed with quantifiably-valued actions.

Entities

Institutions

  • arXiv

Sources