ARTFEED — Contemporary Art Intelligence

Small Reasoning Model with Agentic Workflow Outperforms Larger LLMs in Medical Diagnosis

ai-technology · 2026-07-29

A study published on arXiv (2607.22555) presents the DeepLens Diagnosis Agent, which features a five-stage process integrating a compact medical reasoning model (JSL Medical Small 7B v2) with retrieval-augmented generation (RAG) and a structured workflow. This agent facilitates clinical extraction, methodical retrieval, limited candidate generation, evidence triangulation, and verifiable final decisions. It recorded a top-1 diagnostic accuracy of 60.14% on the 915-case DiagnosisArena benchmark, outperforming other small and medium models. In contrast, the same model without the agent’s workflow achieved only 23.99%, indicating a +36 point improvement solely due to the workflow. The paper contends that single-shot prompting of advanced LLMs often leads to fragile diagnostic reasoning, while the agentic pipeline enhances reliability and transparency.

Key facts

  • DeepLens Diagnosis Agent uses a five-stage pipeline
  • Based on JSL Medical Small 7B v2 model
  • Incorporates retrieval-augmented generation (RAG)
  • Achieved 60.14% top-1 accuracy on DiagnosisArena benchmark
  • Same model without agent workflow scored 23.99%
  • Workflow design yielded +36 point improvement
  • Benchmark had 915 cases
  • Paper published on arXiv with ID 2607.22555

Entities

Institutions

  • arXiv

Sources