Search papers, labs, and topics across Lattice.
Affiliation:
2
0
6
It is found that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth, not only on device bandwidth.
Simulating LLM inference with Kavier reveals how different caching strategies can drastically impact performance, sustainability, and efficiency, offering a crucial tool for optimizing real-world deployments.