LLM Prompting Alone Insufficient for Detecting Shared Decision-Making in Pediatric Encounters
A recent investigation published on arXiv (2608.14792) examines the effectiveness of zero-shot prompting in a large language model (LLM) for identifying shared decision-making (SDM) behaviors during actual pediatric clinical interactions. This study scrutinized 21 audio-recorded outpatient surgical discussions involving 19 distinct patients and 7,566 utterance segments, totaling around 6.1 hours. It compared the zero-shot local LLM (Qwen 2.5 32B) against a supervised classifier utilizing frozen sentence embeddings and a logistic stack. Coders identified 12 SDM behaviors, achieving a human-human macro Cohen's kappa of 0.695. The zero-shot LLM recorded a macro kappa of 0.139 (95% CI 0.111-0.164), while the supervised classifier performed better (exact figure not specified). These results indicate that prompting alone is inadequate for effective SDM detection, emphasizing the importance of robust evaluation and leakage control in clinical communication analysis.
Key facts
- Study analyzed 21 audio-recorded outpatient surgical decision encounters
- 19 unique patients involved
- 7,566 utterance segments analyzed
- Total duration approximately 6.1 hours
- Trained coders labeled 12 SDM behaviors
- Human-human macro Cohen's kappa = 0.695
- Zero-shot LLM (Qwen 2.5 32B) achieved macro kappa = 0.139 (95% CI 0.111-0.164)
- Supervised classifier outperformed zero-shot LLM
- Evaluation used patient-grouped outer folds with inner cross-fitted thresholds and patient-resampled confidence intervals
- Study emphasizes leakage control and nested evaluation
Entities
Institutions
- arXiv