VITA RAG System Outperforms Frontier LLMs on HealthBench
A new system called VITA has been created, and it’s a retrieval-augmented generation (RAG) tool that performs impressively on the HealthBench medical benchmark, either matching or exceeding the latest large language models (LLMs). Tailored for contextual knowledge retrieval in India and similar low- and middle-income countries (LMICs), VITA achieved a score of 51.9% out of the total rubric points on 4,023 English-language questions, hitting 80.5% of the benchmark and claiming first place. This evaluation was carried out by a GPT-4.1 judge. VITA's strength comes from a specialized database featuring disease guidelines and local antimicrobial resistance information. Although its design is proprietary, the benchmark results and scoring are publicly available to highlight the importance of customized medical AI in resource-constrained settings.
Key facts
- VITA is a retrieval-augmented generation (RAG) system for clinical knowledge retrieval in India and other LMIC settings.
- VITA scored 51.9% of possible rubric points on 4,023 English-language HealthBench questions.
- VITA ranked first on the HealthBench benchmark.
- The evaluation used a GPT-4.1 judge.
- VITA retrieves from a curated corpus including disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols.
- The benchmark, physician-written rubrics, and full response and scoring outputs are public for independent verification.
- The results challenge the claim that general-purpose LLMs match or exceed specialized clinical AI tools.
- The comparison is based on a narrow set of systems and benchmarks developed largely in high-income settings.
Entities
Institutions
- VITA
- HealthBench
- GPT-4.1
Locations
- India