Efficient memristor accelerator for transformer self-attention functionality
| dc.contributor.author | Bettayeb, Meriem | |
| dc.contributor.author | Halawani, Yasmin | |
| dc.contributor.author | Khan, Muhammad Umair | |
| dc.contributor.author | ETAL.. | |
| dc.date.accessioned | 2024-10-28T08:25:12Z | |
| dc.date.available | 2024-10-28T08:25:12Z | |
| dc.date.issued | 2024-10-15 | |
| dc.description | Transformer networks combine the advantages of convolutional neural networks (CNN) along with attention models. The attention mechanism is well recognised and very effective owing to its capacity to process input sequences of diverse durations and prioritize the most relevant components of the input. | |
| dc.description.abstract | The adoption of transformer networks has experienced a notable surge in various AI applications. However, the increased computational complexity, stemming primarily from the self-attention mechanism, parallels the manner in which convolution operations constrain the capabilities and speed of convolutional neural networks (CNNs). The self-attention algorithm, specifically the matrix-matrix multiplication (MatMul) operations, demands a substantial amount of memory and computational complexity, thereby restricting the overall performance of the transformer. This paper introduces an efficient hardware accelerator for the transformer network, leveraging memristor-based in-memory computing. The design targets the memory bottleneck associated with MatMul operations in the self-attention process, utilizing approximate analog computation and the highly parallel computations facilitated by the memristor crossbar architecture. Remarkably, this approach resulted in a reduction of approximately 10 times in the number of multiply-accumulate (MAC) operations in transformer networks, while maintaining 95.47% accuracy for the MNIST dataset, as validated by a comprehensive circuit simulator employing NeuroSim 3.0. Simulation outcomes indicate an area utilization of 6895.7 μm2, a latency of 15.52 seconds, an energy consumption of 3 mJ, and a leakage power of 59.55 μW. The methodology outlined in this paper represents a substantial stride towards a hardware-friendly transformer architecture for edge devices, poised to achieve real-time performance. Keywords: Memristor, Accelerator, Operations, Simulation | |
| dc.identifier.citation | Bettayeb, M., Halawani, Y., Khan, M. U., Saleh, H., & Mohammad, B. (2024). Efficient memristor accelerator for transformer self-attention functionality. Scientific Reports, 14(1), 24173. | |
| dc.identifier.doi | https://doi.org/10.1038/s41598-024-75021-z | |
| dc.identifier.uri | https://repository.adu.ac.ae/handle/1/6871 | |
| dc.language.iso | en | |
| dc.publisher | Nature Research | |
| dc.title | Efficient memristor accelerator for transformer self-attention functionality | |
| dc.type | Article |
