Guideline-Based Evaluation of Five Generative Artificial Intelligence Chatbot Platforms for Clinician-Oriented Questions on Open Temporomandibular Joint Surgery

dc.contributor.authorGaş, Selin
dc.contributor.authorSulukan, Erdinç
dc.contributor.authorKorkmaz, Büşra
dc.date.accessioned2026-09-25T12:14:52Z
dc.date.issued2026
dc.departmentDiş Hekimliği Fakültesi
dc.description.abstractObjective: To compare the quality and readability of responses from five generative artificial intelligence chatbot platforms to clinician-oriented questions on open temporomandibular joint (TMJ) surgery against guideline-based reference answers. Material and methods: Forty questions across eight domains were submitted on 16 July 2026 to ChatGPT (GPT5.5), Claude (Opus 4.8), Gemini (3.1 Pro), Grok (4) and Perplexity (Pro), each via its paid tier at default settings (200 responses). Two blinded oral and maxillofacial surgeons applied the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool, the Global Quality Score (GQS) and a five-point overall quality rating using a priori key elements and written anchors. Readability was assessed with the Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES); platforms were compared with the Friedman test and Bonferroni-corrected Wilcoxon post-hoc tests; inter-rater reliability used intraclass correlation coefficients (ICC). Results: Inter-rater reliability was good to excellent (average-measure ICC 0.93−0.99). All outcomes except clarity differed among platforms (P < .001). Perplexity achieved the highest QAMAI total (27.5 ± 1.4; Kendall’s W = 0.75), largely through retrieval-based source provision; excluding this domain, Claude, Perplexity and Gemini converged. Claude had the highest GQS (4.5 ± 0.6), Gemini the highest overall rating (4.3 ± 0.7); Grok scored lowest. Claude and Gemini were least readable (median FKGL 22.7 and 26.1; FRES −11.6 and −7.5). Length did not explain scores within platforms. Conclusions: Performance varied substantially across platforms and dimensions; no platform optimized all outcomes. These findings describe informational quality, not clinical safety or decision-making, and support specialist verification before clinical use.
dc.identifier.doi10.1016/j.jormas.2026.102972
dc.identifier.issn2468-8509
dc.identifier.issn2468-7855
dc.identifier.issue5
dc.identifier.urihttps://hdl.handle.net/11363/12676
dc.identifier.volume127
dc.identifier.wos001876185500001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.institutionauthorGaş, Selin
dc.institutionauthorKorkmaz, Büşra
dc.language.isoen
dc.publisherELSEVIER, RADARWEG 29, 1043 NX AMSTERDAM, NETHERLANDS
dc.relation.ispartofJOURNAL OF STOMATOLOGY ORAL AND MAXILLOFACIAL SURGERY
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.subjectArtificial intelligence
dc.subjectGenerative artificial intelligence
dc.subjectLarge language models
dc.subjectTemporomandibular joint
dc.subjectArthroplasty
dc.subjectClinical decision-making
dc.titleGuideline-Based Evaluation of Five Generative Artificial Intelligence Chatbot Platforms for Clinician-Oriented Questions on Open Temporomandibular Joint Surgery
dc.typeArticle

Dosyalar

Orijinal paket

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
Makale / Article.pdf
Boyut:
513.97 KB
Biçim:
Adobe Portable Document Format

Lisans paketi

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
license.txt
Boyut:
1.17 KB
Biçim:
Item-specific license agreed upon to submission
Açıklama: