Research

Can AI Reliably Answer Patients’ Questions After Orthopedic Surgery? A New Study Puts Leading Large Language Models to the Test


As AI becomes increasingly integrated into healthcare, many hospitals are exploring whether chatbots can help answer patients' questions after surgery. Whether it's about wound care, pain relief, movement restrictions, or warning signs of complications, AI can provide quick answers while reducing the workload of healthcare staff.

But can these systems be trusted when it comes to patient care? Researchers at the Department of Orthopaedic Surgery, Balgrist Unversity Hospital in Zurich put four leading AI models to the test. The goal was to see how accurately, clearly, and safely they could answer common questions from patients recovering after orthopaedic surgery.

The models evaluated were commercial two large language models (GTP-5, Claude 4.5 Sonnet) and two locally deployed alternatives (GTP-OSS and Apertus). 

 

1-1

 

How the Study Was Conducted

To test the AI models, the researchers created 20 realistic scenarios based on questions that patients commonly ask after orthopaedic surgery. These covered recovery after hip, knee, shoulder, and spine procedures and included questions such as "Why is my knee still swollen?" and "When can I start driving again?" (Figure 1).

Each AI model was asked the same set of questions under the same conditions. The models had no access to the internet, external databases, or other tools, ensuring that their answers relied solely on their built-in knowledge.

 

Which AI model provides the best answer? Common postoperative orthopaedic questions were evaluated across four large language models: GPT-5, Claude 4.5 Sonnet, GPT-OSS, and Apertus.

Figure 1: Which AI model provides the best answer? Common postoperative orthopaedic questions were evaluated across four large language models: GPT-5, Claude 4.5 Sonnet, GPT-OSS, and Apertus.

Responses were then independently evaluated by four orthopaedic clinicians using QUEST framework, a practical tool designed to assess the quality of AI-generated responses.

 

Key Findings

The study found that Claude 4.5 Sonnet and GPT-5 provided the highest-quality responses overall, consistently delivering accurate, clear, and trustworthy answers to common postoperative questions. GPT-OSS also performed well but less consistently, while Apertus showed weaker performance, particularly in areas requiring clinical reasoning and patient guidance (Figure 2).

Performance of the four AI models across the evaluation domains of the QUEST framework.

Figure 2: Performance of the four AI models across the evaluation domains of the QUEST framework. Larger areas indicate stronger overall performance. Adapted from Lanter et al. J Exp Orthop. 2026;13:e70813 (CC BY). https://doi.org/10.1002/jeo2.70813. Abbreviations: ACC=Accuracy, COMP=Completeness, REL=Relevance, CLIN=Clinical Appropriateness, EVID=Evidence-Based, CLAR=Clarity, FAB=Fabrication, UND=Understanding, SAT=Satisfaction, SELF=Self-awareness, COMPR=Comprehensiveness.

Safety Matters

One of the study's findings was the difference in safety performance. Apertus generated harmful or fabricated information in approximately 22.5% of evaluations, with harmful content appearing in 9 of 20 responses and fabricated information in 11 of 20 responses. By comparison, GPT-5, Claude 4.5 Sonnet, and GPT-OSS showed very low rates of harmful or fabricated content.

 

The Privacy Trade-Off

While GPT-5 and Claude 4.5 Sonnet performed best, their cloud-based deployment may raise privacy concerns for hospitals. Local AI models provide greater control over patient data, but in this study they were less accurate and reliable.

 

The Future: Specialized Healthcare AI

Researchers suggest that the future may lie in smaller, specialized AI models trained on orthopaedic guidelines, hospital protocols, and medical evidence. Such models could combine the privacy advantages of local deployment with high-quality medical advice.

 

Bottom Line

This study showed that GPT-5 and Claude 4.5 Sonnet delivered the most reliable answers to common postoperative orthopaedic questions. While AI has clear potential to support patient communication, careful testing and strong safety safeguards remain essential before widespread clinical use.

 

To learn more, read the full study here: https://doi.org/10.1002/jeo2.70813

 

Similar posts

Learn Smarter, Stay Updated with AUGMEDI

Stay ahead in medical education and clinical practice. Get the latest updates, tips, and insights delivered straight to your inbox.