As AI becomes increasingly integrated into healthcare, many hospitals are exploring whether chatbots can help answer patients' questions after surgery. Whether it's about wound care, pain relief, movement restrictions, or warning signs of complications, AI can provide quick answers while reducing the workload of healthcare staff.
But can these systems be trusted when it comes to patient care? Researchers at the Department of Orthopaedic Surgery, Balgrist Unversity Hospital in Zurich put four leading AI models to the test. The goal was to see how accurately, clearly, and safely they could answer common questions from patients recovering after orthopaedic surgery.
The models evaluated were commercial two large language models (GTP-5, Claude 4.5 Sonnet) and two locally deployed alternatives (GTP-OSS and Apertus).
To test the AI models, the researchers created 20 realistic scenarios based on questions that patients commonly ask after orthopaedic surgery. These covered recovery after hip, knee, shoulder, and spine procedures and included questions such as "Why is my knee still swollen?" and "When can I start driving again?" (Figure 1).
Each AI model was asked the same set of questions under the same conditions. The models had no access to the internet, external databases, or other tools, ensuring that their answers relied solely on their built-in knowledge.
Responses were then independently evaluated by four orthopaedic clinicians using QUEST framework, a practical tool designed to assess the quality of AI-generated responses.
The study found that Claude 4.5 Sonnet and GPT-5 provided the highest-quality responses overall, consistently delivering accurate, clear, and trustworthy answers to common postoperative questions. GPT-OSS also performed well but less consistently, while Apertus showed weaker performance, particularly in areas requiring clinical reasoning and patient guidance (Figure 2).
One of the study's findings was the difference in safety performance. Apertus generated harmful or fabricated information in approximately 22.5% of evaluations, with harmful content appearing in 9 of 20 responses and fabricated information in 11 of 20 responses. By comparison, GPT-5, Claude 4.5 Sonnet, and GPT-OSS showed very low rates of harmful or fabricated content.
While GPT-5 and Claude 4.5 Sonnet performed best, their cloud-based deployment may raise privacy concerns for hospitals. Local AI models provide greater control over patient data, but in this study they were less accurate and reliable.
Researchers suggest that the future may lie in smaller, specialized AI models trained on orthopaedic guidelines, hospital protocols, and medical evidence. Such models could combine the privacy advantages of local deployment with high-quality medical advice.
This study showed that GPT-5 and Claude 4.5 Sonnet delivered the most reliable answers to common postoperative orthopaedic questions. While AI has clear potential to support patient communication, careful testing and strong safety safeguards remain essential before widespread clinical use.
To learn more, read the full study here: https://doi.org/10.1002/jeo2.70813