LLM Shows Metacognitive Sensitivity in Medical Reasoning Benchmark
A recent investigation published on arXiv (2608.14552) assesses the metacognitive capabilities of large language models (LLMs) in the context of medical reasoning. Researchers created a controlled clinical benchmark inspired by psychophysics to examine diagnostic choices and confidence levels within a medical LLM, specifically targeting the differentiation between probable Alzheimer-type neurocognitive disorder (AT-NCD) and depression-related cognitive impairment (DRCI). The benchmark featured 45 synthetic scenarios with variations in evidence strength, conflicting data, and absent information, presented across three prompt types, resulting in 135 trials. In a preliminary test with gpt-4.1-nano, the model demonstrated 93.5% diagnostic accuracy, an average confidence of 78.4%, and an AUROC2 of 0.876. Confidence was found to rise with evidence clarity and decline when information was lacking, suggesting that LLMs possess metacognitive sensitivity essential for clinical applications. This study emphasizes LLMs' potential in medical diagnostics while highlighting the importance of evaluating their confidence calibration.
Key facts
- Study from arXiv:2608.14552 evaluates LLM metacognitive sensitivity in medical reasoning.
- Benchmark focuses on distinguishing Alzheimer-type neurocognitive disorder (AT-NCD) from depression-related cognitive impairment (DRCI).
- 45 synthetic vignettes varied evidence strength, conflicting evidence, and missing information.
- Each vignette presented under three prompt variants, yielding 135 trials.
- Pilot run used gpt-4.1-nano, producing valid structured outputs for all trials.
- Diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876.
- Confidence increased with evidence distance from diagnostic boundary and decreased with missing information.
- Findings suggest LLMs can exhibit metacognitive sensitivity, important for clinical use.
Entities
Institutions
- arXiv