Search papers, labs, and topics across Lattice.
This survey analyzes recent advancements in deploying Transformer models on FPGA platforms, focusing on operational performance metrics such as throughput and latency, as well as energy efficiency. By comparing FPGA accelerators to traditional CPUs and GPUs, the study highlights the unique advantages of FPGAs, including implementation flexibility and suitability for on-site deployment. The systematic literature review and taxonomy of techniques presented serve as a comprehensive resource for researchers and practitioners looking to optimize Transformer inference on FPGAs.
FPGAs could revolutionize Transformer model deployment by offering superior energy efficiency and latency compared to traditional hardware.
With the rapid and continuous growth in the incorporation of machine learning models based on the Transformer architecture, capable deployment is in high demand. In this context, capable deployment refers to operational performance aspects, e.g., throughput and latency, as well as efficiency aspects, e.g., energy consumption. When it comes to the task of inference using such models, purpose-built hardware accelerators provide a lucrative alternative to common deployment choices, such as Central Processing Units (CPUs) and Graphics Processing Units (GPUs). The Field Programmable Gate Array (FPGA) platforms category is an example of such alternative accelerators, promising implementation flexibility, energy efficiency, improved latency and suitability for on-site deployment. We investigate the most recent advances, trends, and design choices for Transformer inference on FPGA platforms. We perform a systematic literature review, extracting and delving into preferred techniques for implementation and optimisation. This study and the provided taxonomy of topics could act as a guide for researchers from the academia and industry alike.