GraphRareBench: A Benchmark for Rare-Disease Diagnosis
GraphRareBench is a new benchmark for phenotype-driven rare-disease diagnosis, introduced in arXiv:2607.24878. It contains 2,365 ontology-derived cases and 18,093 target-confounder pairs, each with coarsened HPO queries, fixed candidate pools, graph-defined hard confounders, and source-linked evidence records. On a 237-case test split, supervised rankers achieved MRRs of 0.640–0.740 and target-over-confounder accuracies of 0.898–0.916. Agents using Agents-A1 and DeepSeek-V4-Flash achieved MRRs of 0.746 and 0.718, with no significant difference in paired MRR but differences in target-evidence coverage.
Key facts
- GraphRareBench is a provenance-preserving benchmark for rare-disease diagnosis.
- It includes 2,365 ontology-derived cases and 18,093 target-confounder pairs.
- Each case has a coarsened HPO query, fixed candidate pool, graph-defined hard confounders, and source-linked evidence records.
- Supervised rankers achieved MRRs from 0.640 to 0.740 on the test split.
- Target-over-confounder accuracies ranged from 0.898 to 0.916.
- Agents-A1 and DeepSeek-V4-Flash achieved MRRs of 0.746 and 0.718.
- The paired MRR difference between agents was not statistically significant.
- Target-evidence coverage differed between the two agents.
Entities
—