Most medical AI evaluations ask whether a model can answer a benchmark question or generate a plausible differential. We wanted to study something harder and more operationally honest: when our AI proposes a real care action inside a live primary care workflow, does the physician agree?
Today we're publishing our public research at Lotus AI. Across 1,678 AI-suggested care actions for 244 patients (follow-up messages, lab orders, and prescriptions generated by our physician copilot, Lotus Cortex), here's what we found:
→ Among actions that reached a final approve-or-reject decision, physicians approved 98%.
→ Action types behaved very differently. Follow-up messages were the most consistently accepted. Prescriptions carried the highest rejection rate at 10.2%, which is exactly where you want stricter guardrails.
→ Lab orders improved with iteration. More revision cycles between Cortex and the physician correlated with higher approval (Spearman ρ = 0.20, p < 0.001). The physician-agent loop is doing real work.
→ The hardest problem isn't rejection. It's pending. ~30% of proposed actions never reached a final verdict, which is the real operational bottleneck for clinical AI.
A note on why we're sharing this:
Our mission is to build the best clinical AI in the world, and safety has to be load-bearing, not marketing. That means evaluating our systems on what physicians will actually act on, being public about the methodology and its limits, and treating clinician agreement as a first-class signal alongside benchmark performance.
To the engineers, researchers, and physicians thinking seriously about AI in healthcare: we welcome your reactions and critiques. This is the first of many.
Read the full paper: https://lnkd.in/gwqEwrFa