•  
  •  
 

Abstract

Background/purpose: Natural smile photographs are universally accessible yet no automated method achieves clinically acceptable per-tooth detection from these images. Existing frameworks require lip-retracted intraoral views, and prior work on smile images achieved insufficient accuracy. This study developed a per-tooth detection and segmentation pipeline for natural smile photographs and assessed its applicability to digital smile design.

Materials and methods: A total of 1,471 smile photographs were annotated via model-assisted labeling. Fifteen You Only Look Once (YOLO) models across three architecture families and five size tiers were compared. A two-stage pipeline combined the best-performing detector with high-quality Segment Anything Model (HQ-SAM) for per-tooth mask segmentation. Segmentation quality was evaluated against manually drawn ground truth on a 75-image subset, and intra-rater reliability was assessed on 20 images.

Results: YOLO12x achieved the highest mean average precision (mAP50-95) of 0.915. YOLO12 variants occupied the top two positions in precision. HQ-SAM segmentation yielded a mean dice similarity coefficient (Dice) of 0.970 ± 0.055 on 709 matched teeth without dental-specific fine-tuning. A domain comparison confirmed that the smile-trained model substantially outperformed a retracted-view model (mean average precision at intersection over union (IoU) 0.50 [mAP50] 0.959 vs 0.600).

Conclusion: Per-tooth detection and segmentation from natural smile photographs are feasible using data-efficient fine-tuning on a modestly sized, progressively annotated dataset. The narrow performance gap across model sizes suggests compact variants are viable for resource-constrained deployment. Zero-shot HQ-SAM segmentation achieves strong agreement with manual annotations without dental-specific fine-tuning.

Publication Date

2026

Received Date

June 1 2026

Accepted Date

June 17 2026

Final Revision Date

June 15 2026

Share

COinS