

Doctors spend hours on routine tasks: collecting medical history, ordering tests, synthesizing results from various systems. What if an AI agent could take on most of this work? A study published in Nature showed that the autonomous medical AI agent MIRA outperformed doctors in 87.8% of simulated clinical scenarios, working with Electronic Health Record (EHR) data and achieving 88.9% diagnostic accuracy.
In modern medicine, doctors are overwhelmed with administrative work: searching for information in EHRs, reconciling data, ordering tests. This consumes valuable time that could be dedicated to patients and increases the risk of errors due to human factors. Such routine not only reduces efficiency but also leads to specialist burnout. However, there is a way to alleviate this burden using AI agents.
Large Language Models (LLMs) have long demonstrated impressive results in medical exams and are capable of answering complex clinical questions. They can cite thousands of articles, find non-obvious connections, and formulate hypotheses. However, a significant gap existed between this "knowledge" and its real-world application in a clinic's operational processes. Traditional medical AI tools often limited themselves to the role of search engines or text generators, lacking the ability to actively participate in the complex, multi-stage process of clinical decision-making.
True diagnosis and treatment planning are not just about information retrieval. They involve gathering detailed patient history, ordering necessary tests, synthesizing sometimes contradictory data, continuously updating hypotheses, and adjusting the course of treatment. All of this occurs within Electronic Health Record (EHR) systems, which require strict adherence to coding protocols and complex navigation. Doctors spent an enormous amount of time manually processing this data, which slowed down the process and increased the likelihood of errors.
Existing automation solutions in medicine were typically highly specialized. For example, one system might help with X-ray interpretation, another with blood test analysis, and a third with drug information retrieval. But none could integrate these functions into a single, continuous process that mimicked a doctor's thinking and actions. Each solution required manual data entry or switching between interfaces, negating some of the benefits of automation.
The medical community needed not just a tool, but a full-fledged AI agent capable of autonomously navigating the EHR environment, making data-driven decisions, ordering tests, and formulating diagnoses and treatment plans. This meant transitioning from a passive "assistant" to an active "participant" in the process, one that not only provides information but also acts like a doctor, albeit in a controlled environment.
Researchers developed MIRA as an autonomous AI agent capable of operating within isolated EHR environments. The main idea was for MIRA not just to process data, but to "live" within the system, interacting with it autonomously. The agent was equipped with 11 specialized digital tools and had access to over 85,000 operational choices, allowing it to mimic a wide range of clinical actions.
MIRA's functionality included:
A key aspect of the design was to ensure MIRA's autonomy in decision-making while retaining the ability for human oversight and correction in critical situations.
MIRA was not directly implemented into real clinical practice. Instead, to evaluate its effectiveness and safety, a controlled simulation was created using 574 real clinical cases from the MIMIC-IV database. This allowed for a comparison of the AI agent's performance with that of experienced doctors under identical conditions, eliminating the influence of external factors.
Testing phases included:
This approach allowed for a thorough analysis of the agent's performance before its potential application in real-world settings, ensuring maximum safety and reliability.
The study yielded impressive results. MIRA achieved 88.9% diagnostic accuracy across all 574 MIMIC-IV cases. In direct comparison with the group of doctors, MIRA's accuracy was 87.8%, significantly higher than that of experienced human specialists under the same simulated conditions.
| Participant | Average Diagnostic Accuracy |
|---|---|
| MIRA (across all 574 cases) | 88.9% |
| MIRA (compared to doctors, 311 cases) | 87.8% |
| Certified Doctors | 78.1% |
| Mixed-level team (residents + certified doctors) | 71.1% |
MIRA performed particularly well with diagnoses such as appendicitis and pancreatitis, achieving 100% completeness of detection for laparoscopic appendectomies. Importantly, the AI agent did not resort to excessive test ordering; its choices remained below historical baseline levels, indicating high efficiency and cost-effectiveness.
Safety assessments were also encouraging: an independent medical review of 56 patient-level outcomes and 468 prescriptions written by MIRA showed no high-intensity drug interactions, renal dosing incompatibilities, or discrepancies between medications and allergies. The agent also achieved a perfect completeness of detection score (1.00) in critical hospitalization decisions.
The MIRA study results demonstrate the enormous potential of AI agents to transform healthcare. If your clinic faces physician overload with routine tasks, lengthy information retrieval from EHRs, and the need to optimize the diagnostic process, then AI agents could be the solution. Here's where to start:
Despite the impressive results, the study authors emphasize that MIRA and similar AI agents do not replace expert human staff. They require constant human oversight and patient-level safeguards. However, their potential to transform healthcare is immense.
If this case sounds like what's happening in your company, our manager can help: he'll analyze your business and niche for free and point out where an AI agent would bring a real result in your case. Message the manager