LLM Diagnostic Reliability Tested Across Consistency, Manipulation, and Context
A research study assessed the diagnostic accuracy of Google Gemini 2.0 Flash and OpenAI ChatGPT-4o based on three criteria: consistency with rephrased inputs, vulnerability to extraneous prompt content, and ability to adapt to additional clinical context. The researchers created 52 clinical scenarios, adjusting each under controlled settings. To test consistency, scenarios were reworded with variations in demographics, language, and examination details while maintaining the diagnostic essence. Susceptibility was analyzed by incorporating irrelevant yet plausible narrative elements while preserving clinical evidence. For contextual awareness, additional patient history, lifestyle information, or diagnostic results were included to modify the anticipated diagnosis. Physician reviewers evaluated the clinical appropriateness of context-driven adjustments. Both models produced the same diagnoses across all variations and repeated tests.
Key facts
- Study evaluated Google Gemini 2.0 Flash and OpenAI ChatGPT-4o
- 52 clinical scenarios were designed
- Tested consistency, susceptibility to irrelevant content, and contextual awareness
- Both models returned identical diagnoses across all equivalent variants
Entities
Institutions
- OpenAI