- Updated: August 22, 2026
- 1 min read
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease‑specific guidelines, India‑specific antimicrobial resistance data, national formulary constraints, and resource‑limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician‑written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English‑language HealthBench questions (80.5% of the benchmark), scored with a GPT‑4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT‑5.4 (46.1%), o4‑mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%).
Read more about our research and related projects on ubos.tech.

Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.