GuideSkill: Executable LLM Skills for Clinical Reasoning
GuideSkill serves as an external reasoning framework that integrates disease-specific criteria from clinical practice guidelines into actionable functions. It consists of GuideSkill-Zero, which is derived from guidelines, and GuideSkill-Evo, which enhances skills through case-diagnosis pairs. During inference, a large language model (LLM) suggests differential diagnoses, establishes feature relevance, and combines rankings with skill scores. In tests across four benchmarks and four backbones, GuideSkill-Zero enhances macro-average accuracy by 13.45% compared to guideline RAG. Meanwhile, GuideSkill-Evo records the highest macro-average for each backbone, achieving an 18.49% relative improvement over direct inference and raising gold-label skill coverage from 56%.
Key facts
- GuideSkill compiles CPG criteria into executable functions
- GuideSkill-Zero initialized from guidelines
- GuideSkill-Evo refines skills using case-diagnosis pairs
- LLM proposes differential diagnosis at inference
- GuideSkill-Zero improves accuracy by 13.45% over RAG
- GuideSkill-Evo improves over direct inference by 18.49%
- GuideSkill-Evo increases gold-label skill coverage from 56%
- Tested across four benchmarks and four backbones
Entities
—