FuguReport

An AI agent for treatment reasoning over a biomedical tool universe

Authors Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik
Affiliations ATHENA-R1 Evaluation Group / Icahn School of Medicine at Mount Sinai / Mount Sinai Health System / Clalit Research Institute / Harvard Medical School / University of Oxford / Brigham and Women’s Hospital / Broad Institute of MIT and Harvard / Harvard University / Ben Gurion University of the Negev
Categories Application / Biomedical AI / AI agent for drug treatment, Method / Reinforcement Learning / Training on biomedical tool universe, Evaluation / Treatment Effectiveness / Pharmacological treatment accuracy
License CC BY 4.0

Abstract Overview

This paper introduces ATHENA-R1, an AI agent for treatment reasoning that iteratively gathers and integrates evidence from a library of 212 biomedical tools. Rather than relying solely on parametric knowledge, the system identifies missing information, retrieves evidence such as FDA label content, and uses it to update its reasoning step by step. The training pipeline combines self-generated reasoning traces with reinforcement learning using scientific feedback, eliminating the need for human-annotated reasoning demonstrations. The system is comprehensively evaluated on benchmark drug-reasoning tasks, blinded expert reviews for rare diseases, physician-rated hospital cases, and population-scale electronic health record validations.

Novelty

The primary novelty lies in formulating treatment reasoning as a learnable, multi-step evidence-seeking process over a large biomedical tool universe rather than direct answer generation from a language model. Additionally, it introduces a two-stage training scheme utilizing multi-agent self-learning and reinforcement learning with scientific feedback to build reasoning traces and tool-use behaviors without human annotations.

Results

On benchmark evaluations, the agent achieved 94.7% accuracy on DrugPC and 82.9% on TreatmentPC, outperforming models such as GPT-5 and DeepSeek-R1. In blinded expert evaluations, its responses were consistently preferred across clinical criteria, and it successfully reasoned through complex hospitalized-patient cases. Furthermore, retrospective EHR analyses of 5.4 million patients validated its generated adverse-event hypotheses, with predicted associations reaching adjusted odds ratios between 1.48 and 1.84.

Key Points

  1. ATHENA-R1 performs multi-step treatment reasoning by selecting and executing tools from a 212-tool biomedical library in real time to produce evidence-grounded reasoning traces.
  2. The system is trained using a two-level scheme that combines a self-generated instruction dataset with reinforcement learning based on multiaxial scientific feedback.
  3. Empirical evaluation demonstrates state-of-the-art performance across controlled benchmarks, blinded expert assessments, physician reviews of complex cases, and population-scale EHR hypothesis validation.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.