Multi-criteria evaluation of clinical decision-making performance in spinal neurosurgery and physical therapy scenarios: A comparative analysis of artificial intelligence models

dc.contributor.authorTuncer, Cengiz
dc.contributor.authorTekin, Rabia Tugba
dc.contributor.authorUludag, Veysel
dc.contributor.authorKilic, Guven
dc.contributor.authorTaskesen, Ahmet
dc.date.accessioned2026-07-01T11:40:13Z
dc.date.available2026-07-01T11:40:13Z
dc.date.issued2026
dc.departmentDüzce Üniversitesi
dc.description.abstractBackground The integration of AI in healthcare, particularly in clinical decision-making, has shown promising results. This study focuses on evaluating the performance of GPT-4 and GPT-3.5, two advanced AI models, in the context of spinal neurosurgery and physiotherapy, areas that require precise and dynamic decision-making. Methods We conducted a prospective, observational study with 64 participants, including neurosurgeons and physiotherapists, who evaluated AI-generated responses for 10 detailed clinical scenarios. The assessment criteria included diagnostic accuracy, treatment suitability, surgical technique detail, and rehabilitation planning. Each scenario was meticulously crafted to reflect common yet complex clinical situations. Results The study revealed that the GPT-4 consistently outperformed the GPT-3.5 across all the evaluated criteria, with the most significant differences observed in treatment suitability and rehabilitation planning. Statistical analyses, including paired t tests and ANOVA, confirmed the superiority of the GPT-4, highlighting its advanced language processing capabilities and broader medical knowledge base. Reliability analyses further supported these findings. Cronbach's alpha values indicated moderate internal consistency for GPT-4 (alpha = 0.344) and lower consistency for GPT-3.5 (alpha = 0.133). Additionally, Cohen's Kappa values demonstrated moderate agreement for GPT-4 (kappa = 0.65) and fair agreement for GPT-3.5 (kappa = 0.48), further validating the reliability of the participants' evaluations. Conclusions While the GPT-4 has significant potential as a clinical decision support tool, especially in complex and multidisciplinary fields such as spinal neurosurgery and physiotherapy, its recommendations should be carefully integrated with clinical expertise. Further research is essential to enhance its application and ensure that AI can effectively support dynamic medical environments.
dc.identifier.doi10.1007/s00586-026-09795-3
dc.identifier.endpage1108
dc.identifier.issn0940-6719
dc.identifier.issn1432-0932
dc.identifier.issue3
dc.identifier.orcid0000-0003-2400-5546
dc.identifier.orcid0000-0002-3276-5097
dc.identifier.pmid41673318
dc.identifier.scopus2-s2.0-105030018864
dc.identifier.scopusqualityQ1
dc.identifier.startpage1101
dc.identifier.urihttps://doi.org/10.1007/s00586-026-09795-3
dc.identifier.urihttps://hdl.handle.net/20.500.12684/23700
dc.identifier.volume35
dc.identifier.wosWOS:001688021600001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherSpringer
dc.relation.ispartofEuropean Spine Journal
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20260623
dc.subject[Keyword Not Available]
dc.titleMulti-criteria evaluation of clinical decision-making performance in spinal neurosurgery and physical therapy scenarios: A comparative analysis of artificial intelligence models
dc.typeArticle

Dosyalar