Search papers, labs, and topics across Lattice.
FSGen is an innovative framework designed for the efficient generation of AI chip accelerators tailored for large language models (LLMs), addressing the critical challenge of design space exploration (DSE). By integrating fused operator dataflows and leveraging sparsity, FSGen achieves a remarkable 1.4x improvement in power efficiency and up to a 10x speedup while maintaining comparable performance-per-area (PPA) metrics to existing solutions. The framework's advanced PPA estimators significantly enhance design exploration speed and accuracy, resulting in Pareto-optimal designs that outperform previous benchmarks by 58x in figures of merit (FoM).
FSGen achieves a staggering 1.4x power efficiency improvement and 10x speedup for LLM accelerators, redefining the landscape of AI chip design.
With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains. The question of how to efficiently generate optimal AI chip accelerator designs remains unresolved and challenging. Currently, there is a lack of end-to-end design methodologies for efficient design space exploration (DSE). We propose FSGen, an agile framework for attention-based LLM accelerator generation with an early-stage PPA estimator. FSGen supports fused operator dataflows and sparsity with a diverse design space and finds designs with 1.4x better power efficiency or 10x speedup with similar PPA metrics compared to prior work. Pareto-optimal designs have much better performance over a wide range of LLM benchmarks and have 58x better figures of merit (FoM). Design exploration is also faster due to our PPA estimators, which have better accuracy than prior art and reduce DSE runtime drastically.