Search papers, labs, and topics across Lattice.
Duke University, Equal supervision and corresponding authors
3
0
6
ReAct-SQL matches the accuracy of complex text-to-SQL systems while being up to 8 times faster and simpler.
Diffusion LLMs can achieve up to 6.1x higher throughput than autoregressive models by dynamically adjusting decoding granularity based on real-time load, a feat unattainable with fixed-block approaches.
Cut LLM cold starts from minutes to seconds by pre-materializing CUDA graph execution contexts, sidestepping brittle kernel patching and heavyweight checkpointing.