Mamba-caption: Long-range sequence modelling for efficient and accurate image captioning
| dc.contributor.author | Shahzad, Tariq | |
| dc.contributor.author | Aoun, Muhammad | |
| dc.contributor.author | Mazhar, Tehseen | |
| dc.contributor.author | Tariq, Muhammad Usman | |
| dc.date.accessioned | 2026-01-15T06:23:15Z | |
| dc.date.available | 2026-01-15T06:23:15Z | |
| dc.date.issued | 2025-12 | |
| dc.description | At the intersection of natural language processing and computer vision, image captioning was a very popular area of research aimed at generating informative text descriptions for images. | |
| dc.description.abstract | Image captioning has been a problem in vision–language research for a long time. Long-range dependencies and efficiency are challenges for the standard models, such as recurrent neural networks (RNNs) and Transformers. To overcome this, we present Mamba-Caption, an efficient sequence processing model that replaces attention mechanisms with selective state-space modelling. The core novelty is a Mamba-based decoder that substitutes self-attention with selective state-space updates, enabling linear-time caption generation while preserving long-range token dependencies; this decoder is a drop-in language-side component that conditions on a convolutional neural network (CNN) image embedding without domain-specific heuristics. Our model utilizes a CNN encoder, a token embedding layer, and a Mamba-based decoder; the decoder is trained using teacher forcing with a cross-entropy objective. Our model outperforms baselines on all standard metrics when evaluated on the Flickr30k dataset, achieving a Bilingual Evaluation Understudy (BLEU-1) score of 0.83, a Metric for Evaluation of Translation with Explicit ORdering (METEOR) score of 0.79, a Recall-Oriented Understudy for Gisting Evaluation—Longest Common Subsequence (ROUGE-L) score of 0.73, and a Consensus-based Image Description Evaluation (CIDEr) score of 1.30. We further contextualize efficiency via a qualitative/complexity discussion and ablation framing that isolates decoder-side design choices, reinforcing that the gains in efficiency do not sacrifice accuracy. Mamba-Caption can be applied to real-world captioning tasks due to its high efficiency and generalizability. Keywords: CNN, Mamba, CIDEr, Meteor, Blue Score, RNN, Res net, Encoder, Decoder | |
| dc.identifier.citation | Shahzad, T., Aoun, M., Mazhar, T., Tariq, M. U., Ouahada, K., & Hamam, H. (2025). Mamba-caption: Long-range sequence modelling for efficient and accurate image captioning. Array, 100538. | |
| dc.identifier.doi | https://doi.org/10.1016/j.array.2025.100538 | |
| dc.identifier.uri | https://repository.adu.ac.ae/handle/1/7987 | |
| dc.language.iso | en | |
| dc.publisher | Elsevier | |
| dc.title | Mamba-caption: Long-range sequence modelling for efficient and accurate image captioning | |
| dc.type | Article |
Files
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed to upon submission
- Description:
