•  
  •  
 

DOI

https://doi.org/10.1016/j.jds.2025.07.020

First Page

96

Last Page

102

Abstract

Background/purpose Large language models (LLMs), such as GPT-4o and Gemini Advanced, have performed strongly on global medical examinations. However, their capabilities in non-English, dentistry-specific licensing contexts remain unclear. Thus, this study aimed to compare the performance, consistency, and question-generation abilities of GPT-4o and Gemini Advanced in the Korean National Dental Licensing Examination (KNDLE). Materials and methods This study used 1,401 text-based KNDLE questions from 2019 to 2023 in Korean. Each model responded to the questions in three separate runs. Accuracy and consistency were compared with human answers. The models generated new questions in four subject areas and attempted to solve each other's generated items. Paired t-tests and chi-square tests were conducted. Results GPT-4o achieved significantly higher average accuracy than Gemini Advanced (81.1 % vs. 76.6 %, P = 0.013) and showed greater consistency across attempts. Both models performed better in basic sciences than in clinical subjects, such as prosthodontics. In cross-solving tasks, GPT-4o′s performance notably declined in Gemini-generated oral biology questions, indicating interpretation differences. However, the consistency difference between models was not significant (P = 0.578). Conclusion GPT-4o outperformed Gemini Advanced in accuracy, consistency, and alignment with its generated content. However, challenges remain in clinical domains and cross-model understanding, highlighting the potential of LLMs as supportive tools for non-English dental education and question generation while emphasizing the persistent need for expert oversight and domain-specific refinement.

Publication Date

1-1-2026

Share

COinS