Oral tumor detection and localization using detection transformer (DETR) with multi-scale feature extraction and self-supervised learning

dc.contributor.authorKaushik, Pratham
dc.contributor.authorKukreja ,Vinay
dc.contributor.authorJain ,Eshika
dc.contributor.authorAti ,Modafar
dc.contributor.authorHariharan, Shanmugasundaram
dc.date.accessioned2026-07-13T07:45:08Z
dc.date.available2026-07-13T07:45:08Z
dc.date.issued2026-07
dc.descriptionThe condition of oral cancer continues to be one of the primary causes of deaths and diseases related to cancers all around the globe [1] [2]. The annual number of cases of 350,000 new oral cancer patients indicates the importance of early diagnosis to improve healthcare outcomes for patients [3]. Early-stage tumors do not show any symptoms, which leads to a patient only going to medical facilities if their condition is in an advanced stage [4] [5].
dc.description.abstractThis study presents a transformer-based framework for lesion-level detection and localization of oral tumors using the Detection Transformer (DETR) architecture under limited-data medical imaging conditions. The dataset consists of 131 clinically acquired oral images (87 cancerous, 44 non-cancerous). To prevent information leakage and inflated performance estimation, dataset partitioning was performed at the original-image level prior to augmentation, and a 5-fold stratified cross-validation protocol was implemented. Data augmentation was applied exclusively to training subsets, while evaluation was conducted only on non-augmented validation images. To enhance representation learning under low annotation availability, a SimCLR-based self-supervised pretraining stage was incorporated within each cross-validation fold using only training-fold images. The pretrained ResNet-50 backbone was then fine-tuned within the DETR framework for supervised tumor localization. Across cross-validation folds, the proposed model achieved a mean mAP of 0.496 ± 0.027, AP50 of 0.738 ± 0.031, and mean IoU of 0.638 ± 0.019. Bootstrap-based 95% confidence intervals were reported to quantify statistical uncertainty. External validation on an independent oral lesion dataset demonstrated consistent localization performance under domain shift, supporting generalization beyond the primary dataset. Ablation analysis confirmed that self-supervised initialization improves detection stability compared to random initialization. The results demonstrate that a leakage-free transformer-based detection framework combined with self-supervised representation learning can provide reliable lesion-level localization performance in low-resource oral cancer imaging scenarios. keywords Tumor detection ,Detection transformer (DETR), Oral cancer localization, Self-supervised learning (SSL)Attention mechanism, Grad-CAM visualization, Image augmentation
dc.identifier.citationKaushik, P., Kukreja, V., Jain, E., Ati, M., & Hariharan, S. (2026). Oral Tumor Detection and Localization using Detection Transformer (DETR) with Multi-Scale Feature Extraction and Self-Supervised Learning. Array, 100825.
dc.identifier.doihttps://doi.org/10.1016/j.array.2026.100825
dc.identifier.urihttps://repository.adu.ac.ae/handle/1/8386
dc.language.isoen
dc.publisherElesevier
dc.titleOral tumor detection and localization using detection transformer (DETR) with multi-scale feature extraction and self-supervised learning
dc.typeArticle

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
1-s2.0-S2590005626001487-main.pdf
Size:
11.14 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed to upon submission
Description: