How do artificial intelligence chatbots address patient concerns about neuraxial anesthesia?
Medicine Science, cilt.15, sa.3, ss.1171-1176, 2026 (TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 15 Sayı: 3
- Basım Tarihi: 2026
- Doi Numarası: 10.5455/medscience.2026.02.033
- Dergi Adı: Medicine Science
- Derginin Tarandığı İndeksler: TR DİZİN (ULAKBİM)
- Sayfa Sayıları: ss.1171-1176
- Akdeniz Üniversitesi Adresli: Evet
Özet
Artificial intelligence (AI) chatbots are increasingly used by patients seeking medical information before surgery; however, the clinical appropriateness and educational value of chatbot-generated responses regarding neuraxial anesthesia remain uncertain. This study evaluated how three widely accessible AI chatbots—ChatGPT, Gemini, and Copilot—address common patient questions about spinal and epidural anesthesia. Thirty frequently asked questions were developed and grouped into four domains: general information and procedure, safety and risks, intraoperative experience, and postoperative recovery. Each question was submitted identically to the freely accessible, publicly available versions of the three chatbots. Ten board-certified anesthesiologists, blinded to the chatbot source, independently rated each response using a 5-point Likert scale (1 = very inappropriate, 5 = very appropriate). In total, 900 ratings (30 questions × 3 chatbots × 10 evaluators) were analyzed. ChatGPT achieved the highest overall mean score (4.66 ± 0.49), followed by Gemini (4.41 ± 0.49) and Copilot (3.32 ± 0.61). Between-platform differences were statistically significant in a linear mixed-effects model accounting for clustering by rater and question (Wald χ²(2) = 1173.44, p < 0.001). Model-based pairwise comparisons demonstrated that ChatGPT scored higher than Gemini and Copilot, and Gemini scored higher than Copilot (all p < 0.001). Across all four domains, ChatGPT and Gemini consistently outperformed Copilot; ChatGPT ranked highest in three domains, while Gemini showed slightly better performance for intraoperative experience–related questions. In this expert-rating study conducted at a single time point, ChatGPT-generated responses were rated as more clinically appropriate than those produced by Gemini and Copilot. These findings suggest that AI chatbots may support preoperative patient education, but their outputs require cautious interpretation, ongoing expert oversight, and periodic re-evaluation as platforms evolve.