Comparative performance of AI chatbots in dental implantology: insights and limitations
| dc.authorid | https://orcid.org/0009-0004-4149-2061 | |
| dc.authorid | https://orcid.org/0000-0002-3337-9403 | |
| dc.authorid | https://orcid.org/0000-0001-9321-1031 | |
| dc.contributor.author | Uçar, Sultan Merve | |
| dc.contributor.author | Gaş, Selin | |
| dc.contributor.author | Sasany, Rafat | |
| dc.date.accessioned | 2026-07-24T14:08:10Z | |
| dc.date.issued | 2026 | |
| dc.department | Diş Hekimliği Fakültesi | |
| dc.description.abstract | Objective: This study critically evaluated the performance, accuracy, and clinical relevance of three large language models ChatGPT-4o, Claude 3.5, and Gemini 1.5 Pro when answering expert-generated questions on zygomatic implantology. The goal was to determine the extent to which such tools may function as educational or clinical decision supports in maxillofacial surgery. Methods: Thirty-eight standardized questions were developed by four oral and maxillofacial surgeons with advanced expertise in zygomatic implantology. Each model's responses were independently assessed by five calibrated clinical raters using validated metrics DISCERN, GQS, and a 5-point Accuracy Rubric to judge reliability, quality, and factual correctness. Non-parametric statistics (Kruskal-Wallis with Bonferroni post hoc correction; Spearman correlation) were used, and inter-rater reliability was quantified by ICC(2,1) = 0.86-0.91 (p < 0.001). Results: Gemini 1.5 Pro achieved slightly higher mean scores for response quality and accuracy, whereas Claude 3.5 and ChatGPT-4o performed comparably. However, absolute differences were modest (≤ 0.5 points on 5-point scales), indicating relative trends rather than decisive superiority. All models produced readable, clinically relevant content, though variability persisted in the depth and specificity of clinical guidance. Conclusion: Current AI language models exhibit moderate but inconsistent competency when addressing complex implantology scenarios. While Gemini 1.5 Pro scored marginally higher, these differences are unlikely to be of major practical consequence. Continuous validation, transparent reporting of model versions, and expert supervision remain essential before integrating such systems into routine dental education or clinical decision-making. | |
| dc.identifier.citation | Uçar SM, Gaş S, Sasany R. Comparative performance of AI chatbots in dental implantology: insights and limitations. BMC Oral Health. 2025 Dec 17;26(1):147. doi: 10.1186/s12903-025-07426-9. PMID: 41408269; PMCID: PMC12829278. | |
| dc.identifier.doi | 10.1186/s12903-025-07426-9 | |
| dc.identifier.issn | 1472-6831 | |
| dc.identifier.pmid | 41408269 | |
| dc.identifier.uri | https://hdl.handle.net/11363/11931 | |
| dc.indekslendigikaynak | PubMed | |
| dc.institutionauthor | Uçar, Sultan Merve | |
| dc.institutionauthor | Gaş, Selin | |
| dc.institutionauthorid | https://orcid.org/0009-0004-4149-2061 | |
| dc.institutionauthorid | https://orcid.org/0000-0002-3337-9403 | |
| dc.language.iso | en | |
| dc.publisher | BioMed Central Ltd | |
| dc.relation.ispartof | BMC Oral Health | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.subject | Artificial intelligence | |
| dc.subject | Dental implantology | |
| dc.subject | Dental education | |
| dc.subject | Clinical decision-making | |
| dc.title | Comparative performance of AI chatbots in dental implantology: insights and limitations | |
| dc.type | Article |










