ViT-Stain: Vision transformer-driven virtual staining for skin histopathology via global contextual learning

dc.contributor.authorAsaf, Muhammad Zeeshan
dc.contributor.authorGilani, Syed Omer
dc.contributor.authorJavaid, Amber
dc.contributor.authorETAL..
dc.date.accessioned2026-07-20T07:41:26Z
dc.date.available2026-07-20T07:41:26Z
dc.date.issued2026
dc.description.abstractCurrent virtual staining approaches for histopathology slides use convolutional neural networks (CNNs) and generative adversarial networks (GANs). These approaches rely on local receptive fields, struggle to capture global context, and long-range tissue dependencies. This limitation can introduce artifacts in fine textures and cause loss of subtle morphological details. We propose a novel vision transformer-driven virtual staining framework (ViT-Stain) that translates unstained skin tissue images into hematoxylin and eosin (H&E)-equivalent images. The transformer’s self-attention enables ViT-Stain to capture long-range dependencies, preserve global context, and maintain fine textures. We trained ViT-Stain on the E-Staining DermaRepo dataset, which pairs unstained and H&E-stained whole-slide images (WSIs). We validated our model using metrics including SSIM, PSNR, FID, KID, LPIPS, and a novel histology-specific fidelity index (HSFI). Three board-certified pathologists provided feedback for qualitative evaluations. ViT-Stain outperforms leading CNN and GAN models, including Pix2Pix, CycleGAN, CUTGAN, and DCLGAN. It achieves an overall diagnostic concordance of 85% with virtual H&E-stains (Fleiss’ κ=0.88). However, the model requires longer training (about 93 hours on A100 GPUs) and inference times (about 2.9 minutes). Our work advances AI-driven diagnostic reproducibility for high-fidelity clinical settings and aligns with the World Health Organization (WHO) global health goals. Keywords:Hematoxylin, Humans, Image Processing, Computer-Assisted, Neural Networks, Computer, Skin, Staining and Labeling
dc.identifier.citationHussain, M. A., Waris, M. A., Akram, M. U., Khan, M. J., Asaf, M. Z., Javaid, A., ... & Hazzazi, F. (2026). ViT-Stain: Vision transformer-driven virtual staining for skin histopathology via global contextual learning. PloS one, 21(2), e0341311.
dc.identifier.doihttps://doi.org/10.1371/journal.pone.0341311
dc.identifier.urihttps://repository.adu.ac.ae/handle/1/8428
dc.language.isoen
dc.publisherPublic Library of Science
dc.titleViT-Stain: Vision transformer-driven virtual staining for skin histopathology via global contextual learning
dc.typeArticle

Files

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed to upon submission
Description: