Comparative performance of AI chatbots in dental implantology: insights and limitations

dc.authoridhttps://orcid.org/0009-0004-4149-2061
dc.authoridhttps://orcid.org/0000-0002-3337-9403
dc.authoridhttps://orcid.org/0000-0001-9321-1031
dc.contributor.authorUçar, Sultan Merve
dc.contributor.authorGaş, Selin
dc.contributor.authorSasany, Rafat
dc.date.accessioned2026-07-24T14:08:10Z
dc.date.issued2026
dc.departmentDiş Hekimliği Fakültesi
dc.description.abstractObjective: This study critically evaluated the performance, accuracy, and clinical relevance of three large language models ChatGPT-4o, Claude 3.5, and Gemini 1.5 Pro when answering expert-generated questions on zygomatic implantology. The goal was to determine the extent to which such tools may function as educational or clinical decision supports in maxillofacial surgery. Methods: Thirty-eight standardized questions were developed by four oral and maxillofacial surgeons with advanced expertise in zygomatic implantology. Each model's responses were independently assessed by five calibrated clinical raters using validated metrics DISCERN, GQS, and a 5-point Accuracy Rubric to judge reliability, quality, and factual correctness. Non-parametric statistics (Kruskal-Wallis with Bonferroni post hoc correction; Spearman correlation) were used, and inter-rater reliability was quantified by ICC(2,1) = 0.86-0.91 (p < 0.001). Results: Gemini 1.5 Pro achieved slightly higher mean scores for response quality and accuracy, whereas Claude 3.5 and ChatGPT-4o performed comparably. However, absolute differences were modest (≤ 0.5 points on 5-point scales), indicating relative trends rather than decisive superiority. All models produced readable, clinically relevant content, though variability persisted in the depth and specificity of clinical guidance. Conclusion: Current AI language models exhibit moderate but inconsistent competency when addressing complex implantology scenarios. While Gemini 1.5 Pro scored marginally higher, these differences are unlikely to be of major practical consequence. Continuous validation, transparent reporting of model versions, and expert supervision remain essential before integrating such systems into routine dental education or clinical decision-making.
dc.identifier.citationUçar SM, Gaş S, Sasany R. Comparative performance of AI chatbots in dental implantology: insights and limitations. BMC Oral Health. 2025 Dec 17;26(1):147. doi: 10.1186/s12903-025-07426-9. PMID: 41408269; PMCID: PMC12829278.
dc.identifier.doi10.1186/s12903-025-07426-9
dc.identifier.issn1472-6831
dc.identifier.pmid41408269
dc.identifier.urihttps://hdl.handle.net/11363/11931
dc.indekslendigikaynakPubMed
dc.institutionauthorUçar, Sultan Merve
dc.institutionauthorGaş, Selin
dc.institutionauthoridhttps://orcid.org/0009-0004-4149-2061
dc.institutionauthoridhttps://orcid.org/0000-0002-3337-9403
dc.language.isoen
dc.publisherBioMed Central Ltd
dc.relation.ispartofBMC Oral Health
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.subjectArtificial intelligence
dc.subjectDental implantology
dc.subjectDental education
dc.subjectClinical decision-making
dc.titleComparative performance of AI chatbots in dental implantology: insights and limitations
dc.typeArticle

Dosyalar

Orijinal paket

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
Makale / Article.pdf
Boyut:
3.03 MB
Biçim:
Adobe Portable Document Format

Lisans paketi

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
license.txt
Boyut:
1.17 KB
Biçim:
Item-specific license agreed upon to submission
Açıklama: