Large language models are becoming remarkably good at medicine. They can pass medical examinations, solve complex clinical vignettes and explain difficult concepts within seconds. For medical students and residents, this creates something previous generations never had: almost unlimited access to an intelligent assistant.
At first glance, this seems like an educational breakthrough. But there is a more important question:
Does having access to better answers actually make us better doctors?
A growing body of research suggests that the answer is more complicated. AI can improve performance. But access to AI does not automatically produce understanding, clinical reasoning or the ability to recognise when an answer should be trusted. That distinction may become one of the central challenges of medical education in the AI era.
The capabilities of modern large language models should not be underestimated. A recent systematic review in BMC Medical Education included 43 studies with approximately 12,400 participants across 19 countries. GPT-4 achieved passing scores on USMLE Steps 1-3 in every included study in which it was tested, with a reported mean accuracy of 86%.
But performance declined as the tasks became more complex. Across the studies included in the review, accuracy dropped to 61-71% for multi-step clinical vignettes and 48-63% for image-dependent questions.
That difference matters. A multiple-choice examination usually provides the information required to solve a problem. Clinical medicine does not. Physicians must determine which questions to ask, what to examine, which investigations are necessary, how to interpret incomplete information and when sufficient evidence exists to act.
High benchmark performance therefore does not automatically translate into clinical competence.
Performance ranges reported in a 2026 systematic review of 43 studies, together with emerging evidence on human–AI interaction in medical education.
There is another important finding in the same review. While students generally perceived LLMs positively, evidence that their use actually improves learning was considerably weaker. Only six of twelve studies examining knowledge acquisition reported modest short-term benefits, and durable retention had not been demonstrated. The authors also rated 86% of the included studies as having a moderate-to-high risk of bias.
This exposes an important distinction:
Performing a medical task and teaching someone to perform that task are not the same problem.
Imagine a medical student working through a diagnostic case. The student develops a differential diagnosis. An AI system then suggests another possibility. Should the student change the diagnosis?
A 2026 study involving 185 pre-clerkship medical students examined exactly this interaction. Students ranked their differential diagnoses before and after receiving access to an AI chatbot.
Overall, AI use significantly improved diagnostic accuracy, with the largest benefit occurring among students who performed poorly before using AI. That is encouraging. But the same experiment revealed a much more interesting problem.
Among the 185 students, diagnostic accuracy improved in 65 after AI interaction but worsened in 26. Both groups nevertheless rated the AI as having significantly influenced their reasoning. The authors concluded that students did not reliably distinguish helpful from harmful AI influence, consistent with automation bias.
This may be more important for medical education than whether an AI model achieves 80%, 90% or 95% on a benchmark. The critical skill in an AI-supported medical environment is not simply knowing how to ask AI a question. It is knowing when to trust the answer.
The obvious solution might seem straightforward: tell learners that AI can make mistakes. That approach has now been tested. Kıyak and colleagues randomised 186 fourth-year medical students to receive diagnostic AI advice either with or without the warning:
WARNING SHOWN TO STUDENTS
“ChatGPT can make mistakes. Check important info.”
The warning did not meaningfully change behaviour. Students changed their diagnosis in 15.3% of responses without the warning and 15.9% with the warning. The difference was not statistically significant (OR 1.09; 95% CI 0.46-2.59; p=0.84).
Importantly, the students were not simply following AI blindly. The authors found that students tended to underweight the AI’s diagnostic advice. The lesson is therefore subtler than “students trust AI too much.”
The problem is calibrated trust. Knowing that AI can be wrong does not tell a learner when it is wrong.
Ideally, an AI system would compensate for this by communicating its own uncertainty accurately. Current systems still struggle with that.
Griot and colleagues evaluated twelve large language models using MetaMedQA, a benchmark designed specifically to assess medical metacognition: whether models recognise the limits of their own knowledge.
Despite strong conventional question-answering performance, the models showed substantial deficiencies in recognising those limits. They could provide confident answers even when the correct answer was absent from the available options.
That creates an important human-AI problem: the learner may not know when the AI is wrong, while the AI may not reliably know when it is wrong either. In medicine, where uncertainty is inherent and decisions can have real consequences, this limitation matters.
Perhaps the clearest demonstration of the difference between AI capability and human capability comes from a randomised clinical trial published in JAMA Network Open.
Physicians were randomised to solve clinical cases either using conventional resources or with additional access to an LLM. The median diagnostic reasoning score was 76% with LLM access versus 74% with conventional resources, a difference that was not statistically significant (p=0.60).
Yet the LLM working alone significantly outperformed the conventional-resources physician group by 16 percentage points.
Think about the implication. The technology itself was highly capable. But giving that technology to a physician did not automatically transfer its full capability to the physician.
That is the key educational problem:
AI capability is not the same as human capability.
A general-purpose AI system is designed to help produce an answer. Medical education has a fundamentally different objective. Its goal is to change what the learner can understand, retrieve and apply.
A future surgeon must recognise anatomical structures, a resident must build a differential diagnosis, a student must retrieve relevant knowledge under pressure, and a physician must recognise when new information contradicts an initial hypothesis.
All of them need mental models strong enough to critically evaluate information generated by both humans and machines.
That requires more than access to information. It requires a learning process:
Structured knowledge → understanding → retrieval → application → feedback → repetition
There is substantial evidence supporting this approach. For example, a prospective randomised cohort study involving 26,258 family physicians and residents evaluated spaced repetition in real-world continuing medical education.
Participants receiving spaced repetition performed better than controls both on later learning assessments (58.03% vs. 43.20%) and on subsequent knowledge-transfer questions (58.33% vs. 52.39%). Repeating material twice produced additional improvements compared with a single repetition.
The educational question is therefore changing.
For generations, one of the central problems was:
How do we give learners access to knowledge?
AI is rapidly reducing that problem. The new question is:
How do we turn unlimited access to information into reliable medical competence?
This distinction is central to how we think about medical education at AUGMEDI.
The goal should not be to compete with general-purpose AI by trying to answer more questions. AI systems are already extraordinarily good at doing that.
Instead, AI becomes considerably more useful when it operates inside a structured medical learning architecture.
For an orthopaedic learner, that progression might look like this:
AI can support every stage of this process. But the learning architecture has to come first.
A general-purpose chatbot starts with a question:
“What would you like to know?”
A structured educational system starts somewhere else:
“What do you need to know next?”
That distinction is fundamental. Learners do not always know what they do not know. Expertise is not built by accumulating isolated answers. It develops when anatomy, pathology, clinical reasoning and procedural knowledge are repeatedly connected, retrieved and applied.
This becomes particularly important in procedural specialties such as orthopaedics. Knowing the name of a nerve is not the same as locating it in three-dimensional anatomy. Locating it in anatomy is not the same as understanding where it appears during a surgical approach. Recognising it during an operation requires another level of expertise again. AI can explain each of those steps. Medical education must connect them.
AI will become an integral part of medical education. The evidence already suggests that it can improve performance, particularly for learners who initially struggle.
The appropriate conclusion is therefore not that medical students need less AI. They need better educational systems around AI:
The defining challenge of medical education is changing. For generations, the problem was access to knowledge. Today, an enormous amount of medical information is only seconds away.
The problem is no longer access to answers. The problem is knowing which answers to trust, what to learn next, and whether you have actually learned it. That is where the next generation of medical education begins.
Bodke P, Jumle R. Large language models in undergraduate medical education: a systematic review of examination performance, knowledge acquisition, and learner perceptions. BMC Medical Education. Published July 30, 2026.
Hayden R, Blanco M, Ramesh S, et al. Diagnostic reasoning with and without AI: automation bias in pre-clerkship medical students. BMC Medical Education. Published July 21, 2026.
Kıyak YS, Coşkun Ö, Budakoğlu Iİ. “ChatGPT can make mistakes” warnings fail: A randomized controlled trial. Medical Education. 2026;60(2):138-142. DOI: 10.1111/medu.70056.
Goh E, Gallo R, Hom J, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open. 2024;7.
Griot M, Hemptinne C, Vanderdonckt J, Yuksel D. Large Language Models lack essential metacognition for reliable medical reasoning. Nature Communications. 2025;16:642.
Price DW, Wang T, O’Neill TR, et al. The Effect of Spaced Repetition on Learning and Knowledge Transfer in a Large Cohort of Practicing Physicians. Academic Medicine. 2025;100(1):94-102. DOI: 10.1097/ACM.0000000000005856.