Efficient memristor accelerator for transformer self-attention functionality

dc.contributor.authorBettayeb, Meriem
dc.contributor.authorHalawani, Yasmin
dc.contributor.authorKhan, Muhammad Umair
dc.contributor.authorETAL..
dc.date.accessioned2024-10-28T08:25:12Z
dc.date.available2024-10-28T08:25:12Z
dc.date.issued2024-10-15
dc.descriptionTransformer networks combine the advantages of convolutional neural networks (CNN) along with attention models. The attention mechanism is well recognised and very effective owing to its capacity to process input sequences of diverse durations and prioritize the most relevant components of the input.
dc.description.abstractThe adoption of transformer networks has experienced a notable surge in various AI applications. However, the increased computational complexity, stemming primarily from the self-attention mechanism, parallels the manner in which convolution operations constrain the capabilities and speed of convolutional neural networks (CNNs). The self-attention algorithm, specifically the matrix-matrix multiplication (MatMul) operations, demands a substantial amount of memory and computational complexity, thereby restricting the overall performance of the transformer. This paper introduces an efficient hardware accelerator for the transformer network, leveraging memristor-based in-memory computing. The design targets the memory bottleneck associated with MatMul operations in the self-attention process, utilizing approximate analog computation and the highly parallel computations facilitated by the memristor crossbar architecture. Remarkably, this approach resulted in a reduction of approximately 10 times in the number of multiply-accumulate (MAC) operations in transformer networks, while maintaining 95.47% accuracy for the MNIST dataset, as validated by a comprehensive circuit simulator employing NeuroSim 3.0. Simulation outcomes indicate an area utilization of 6895.7 μm2, a latency of 15.52 seconds, an energy consumption of 3 mJ, and a leakage power of 59.55 μW. The methodology outlined in this paper represents a substantial stride towards a hardware-friendly transformer architecture for edge devices, poised to achieve real-time performance. Keywords: Memristor, Accelerator, Operations, Simulation
dc.identifier.citationBettayeb, M., Halawani, Y., Khan, M. U., Saleh, H., & Mohammad, B. (2024). Efficient memristor accelerator for transformer self-attention functionality. Scientific Reports, 14(1), 24173.
dc.identifier.doihttps://doi.org/10.1038/s41598-024-75021-z
dc.identifier.urihttps://repository.adu.ac.ae/handle/1/6871
dc.language.isoen
dc.publisherNature Research
dc.titleEfficient memristor accelerator for transformer self-attention functionality
dc.typeArticle

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Efficient memristor accelerator.pdf
Size:
3.84 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed to upon submission
Description: