Mamba-caption: Long-range sequence modelling for efficient and accurate image captioning

dc.contributor.authorShahzad, Tariq
dc.contributor.authorAoun, Muhammad
dc.contributor.authorMazhar, Tehseen
dc.contributor.authorTariq, Muhammad Usman
dc.date.accessioned2026-01-15T06:23:15Z
dc.date.available2026-01-15T06:23:15Z
dc.date.issued2025-12
dc.descriptionAt the intersection of natural language processing and computer vision, image captioning was a very popular area of research aimed at generating informative text descriptions for images.
dc.description.abstractImage captioning has been a problem in vision–language research for a long time. Long-range dependencies and efficiency are challenges for the standard models, such as recurrent neural networks (RNNs) and Transformers. To overcome this, we present Mamba-Caption, an efficient sequence processing model that replaces attention mechanisms with selective state-space modelling. The core novelty is a Mamba-based decoder that substitutes self-attention with selective state-space updates, enabling linear-time caption generation while preserving long-range token dependencies; this decoder is a drop-in language-side component that conditions on a convolutional neural network (CNN) image embedding without domain-specific heuristics. Our model utilizes a CNN encoder, a token embedding layer, and a Mamba-based decoder; the decoder is trained using teacher forcing with a cross-entropy objective. Our model outperforms baselines on all standard metrics when evaluated on the Flickr30k dataset, achieving a Bilingual Evaluation Understudy (BLEU-1) score of 0.83, a Metric for Evaluation of Translation with Explicit ORdering (METEOR) score of 0.79, a Recall-Oriented Understudy for Gisting Evaluation—Longest Common Subsequence (ROUGE-L) score of 0.73, and a Consensus-based Image Description Evaluation (CIDEr) score of 1.30. We further contextualize efficiency via a qualitative/complexity discussion and ablation framing that isolates decoder-side design choices, reinforcing that the gains in efficiency do not sacrifice accuracy. Mamba-Caption can be applied to real-world captioning tasks due to its high efficiency and generalizability. Keywords: CNN, Mamba, CIDEr, Meteor, Blue Score, RNN, Res net, Encoder, Decoder
dc.identifier.citationShahzad, T., Aoun, M., Mazhar, T., Tariq, M. U., Ouahada, K., & Hamam, H. (2025). Mamba-caption: Long-range sequence modelling for efficient and accurate image captioning. Array, 100538.
dc.identifier.doihttps://doi.org/10.1016/j.array.2025.100538
dc.identifier.urihttps://repository.adu.ac.ae/handle/1/7987
dc.language.isoen
dc.publisherElsevier
dc.titleMamba-caption: Long-range sequence modelling for efficient and accurate image captioning
dc.typeArticle

Files

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed to upon submission
Description: