•  
  •  
 

Abstract

Background/purpose: Large language models (LLMs) offer promising education and clinical decision support to help address the global shortage of maxillofacial prosthodontists. To overcome the lack of source traceability in base LLMs, this study evaluated a retrieval-augmented generation (RAG) architecture for answering complex clinical questions.

Materials and methods: An expert panel consisting of five maxillofacial prosthodontists selected 15 high-complexity clinical questions. Responses were generated using base GPT-4o and a naive RAG-GPT-4o model built with domain-specific literatures. Outputs were evaluated across five parameters using a 5-point Likert scale and analyzed via Mann-Whitney U tests.

Results: No statistically significant differences were observed between GPT-4o and RAG-GPT-4o across all parameters (P  > 0.05). The median scores (IQR) for GPT-4o and RAG-GPT-4o were as follows: relevance (4.40 ± 0.40 vs. 4.40 ± 0.50, P = 0.815), clarity (4.40 ± 0.40 vs. 4.20 ± 0.60, P = 0.232), depth (4.20 ± 0.50 vs. 4.00 ± 0.50, P = 0.378), focus (4.40 ± 0.60 vs. 4.00 ± 0.50, P = 0.258), and coherence (4.40 ± 0.30 vs. 4.20 ± 0.50, P = 0.234).

Conclusion: This study demonstrates that GPT-4o possesses significant latent expertise in maxillofacial prosthodontics for clinical questions. While the base LLM and a naive RAG architecture performed comparably across predefined clinical parameters, RAG integration provided an essential advantage: verifiable factual grounding. By anchoring AI outputs in transparent, evidence-based literature, RAG addresses the critical need for clinical safety and clinician trust. Ultimately, while base LLMs offer substantial clinical reasoning capabilities, pairing them with RAG architectures remains necessary to ensure the accountability required for safe integration into specialized dental care, establishing a foundation for future development of  more advanced RAG systems

Publication Date

2026

Share

COinS