LLMs vs Official Crash Coding: Benchmarking Frontier Models on Arkansas Fatal-Crash Narratives
A new study looked at six sophisticated large language models (LLMs) to see how well they could match official crash coding used in Arkansas police reports. Published on arXiv (2607.29064), the research aimed to improve structured crash databases, which are often a hassle to analyze manually. It examined 5,587 narratives from fatal crashes and 5,889 structured records from 2015 to 2025, finding 4,194 that matched. Each LLM received the same prompt to classify six crash attributes, including crash manner and light conditions. The researchers used various metrics like agreement rates and F1 scores to measure accuracy, comparing the results to standard methods. This study could potentially automate crash data coding, aiding traffic safety efforts.
Key facts
- Study benchmarked six frontier LLMs for crash coding from police narratives.
- Data from Arkansas fatal-crash database (2015-2025).
- Linked 5,587 narratives with 5,889 structured records, yielding 4,194 matches.
- Evaluated six crash attributes: manner, non-motorist relation, intersection type, work-zone relation, surface condition, light condition.
- Used identical zero-shot prompt for all LLMs.
- Performance metrics: agreement, macro-F1, Cohen's kappa, coverage, selective agreement.
- Compared against always-majority, always-Unknown, and keyword-rule baselines.
- Study available on arXiv with ID 2607.29064.
Entities
Institutions
- arXiv
Locations
- Arkansas
- United States