FitAQA: New Benchmark Evaluates Multimodal LLMs on Fitness Action Quality
FitAQA has been launched by researchers as a structured benchmark to assess the performance of Multimodal Large Language Models (MLLMs) in evaluating the quality of fitness actions. This benchmark, outlined in a paper available on arXiv (2608.08736), focuses on the relatively neglected field of MLLMs in smart sports training. Current benchmarks typically depend on action-specific annotations and mainly concentrate on final evaluation results, which limits their insights into exercise quality assessments. FitAQA comprises 2,219 videos and 5,512 QA instances covering 30 bodyweight exercises. Collaborating with sports science specialists, it presents a unified taxonomy of 38 common form errors categorized into six quality dimensions: alignment, symmetry, stability, coordination, tempo, and completeness. Additionally, FitAQA establishes three evaluation tasks, including perception for identifying key visual evidence, and two other tasks that are not fully detailed in the abstract. This benchmark seeks to enhance the understanding of MLLMs in fitness action quality assessment, potentially benefiting intelligent sports training innovations.
Key facts
- FitAQA is a new benchmark for evaluating MLLMs in fitness action quality assessment.
- It contains 2,219 videos and 5,512 QA instances across 30 bodyweight exercises.
- The benchmark was developed in collaboration with sports science experts.
- It introduces a unified form error taxonomy with 38 recurring form errors.
- The taxonomy covers six quality dimensions: alignment, symmetry, stability, coordination, tempo, and completeness.
- FitAQA formulates three evaluation tasks, including perception for recognizing visual evidence.
- The paper is available on arXiv with identifier 2608.08736.
- Existing benchmarks lack insight into how models assess exercise quality.
Entities
Institutions
- arXiv