Search papers, labs, and topics across Lattice.
2
0
4
0
Achieving up to 2.3x throughput gains, HYDRA reveals that co-designing architecture and runtime policies is essential for optimizing hybrid LLM workloads on chiplet systems.
Hybrid Mamba-Transformer models can get 4x faster time to first token and 1.4x higher throughput by disaggregating prefill and decode phases onto specialized accelerator packages.