Preformulation Gap: LLM Medical Consultations Fail to Probe Vague Patient Concerns
A recent preprint available on arXiv (identifier 2608.17330v1) investigates the timing of evaluations for large language models during medical consultations, pointing out a 'preformulation gap' where assessments take place after the clinical issue is defined, which differs from actual first-contact interactions. The research analyzed three unnamed API-based models using four vignettes created by physicians, contrasting a baseline scenario with one featuring 'entry-to-care' instructions. The primary protocol generated 24 fixed-script transcripts, while two vignettes used adaptive simulations, contributing an additional 12. In the baseline scenario, models provided self-care recommendations in 9 out of 12 instances, whereas the instruction condition yielded none. The findings advocate for assessing the preformulation gap via observable first-contact behaviors instead of final results.
Key facts
- The preprint is arXiv:2608.17330v1.
- It addresses the 'preformulation gap' in LLM medical consultation.
- Three API-based LLMs were evaluated.
- Four physician-authored multi-turn vignettes were used.
- 24 fixed-script transcripts came from baseline and instruction conditions.
- 12 transcripts came from adaptive standardized-patient simulation in two cases.
- Baseline models gave self-care advice before any patient answer in 9 of 12 cells; instruction cells had 0.
- Structured handoff summaries appeared in 0 of 12 baseline cells and 10 of 12 instruction cells.
Entities
Institutions
- arXiv