Search papers, labs, and topics across Lattice.
5
0
5
3
PLoRA slashes decode latency for multi-LoRA serving by over 6x while using pooled memory and near-data processing, revolutionizing how we deploy specialized AI models.
Transforming context ahead of time can slash time-to-first-token by nearly 12x, revolutionizing LLM agent efficiency.
FlashCP achieves up to 1.63x faster training for large language models by eliminating redundant communication and optimizing workload balance.
Forget GPU-centric designs: AMMA slashes attention latency by 15x and energy consumption by 7x with a memory-centric architecture for long-context LLMs.
LLMs still have a long way to go in AI-aided chip design, with even the best models achieving surprisingly low scores on the new ChipBench benchmark for Verilog generation and reference model creation.