Search papers, labs, and topics across Lattice.
The Hong Kong University of Science and Technology (Guangzhou)
11
1
17
17
Automated strategy optimization using LLMs can boost Sharpe ratios by over 199%, revolutionizing quantitative trading practices.
Xema achieves a staggering 3.7x improvement in service level objectives for diffusion models while slashing planning time from hours to mere minutes.
KernelFlume slashes operational costs by up to 61% while maintaining low latency for long-context LLM decoding, revolutionizing how we scale model serving.
DigenRL achieves up to 2.10x throughput gains for diffusion-based generative LLMs by disaggregating resource allocation and optimizing execution pipelines.
LLM training bottlenecks? ZipCCL achieves up to 1.18x end-to-end speedups by losslessly compressing communication collectives, without sacrificing model quality.
LLMs can leapfrog state-of-the-art scientific algorithms and human-designed solutions, but only if you scale the evaluation loop, not just the model.
Lossless compression can actually *speed up* LLM inference on GPUs, not just shrink model size, thanks to ZipServ's hardware-aware design.
LLM-powered recommendation agents, despite their reasoning prowess, are easily manipulated by contextual biases in high-stakes scenarios like paper review and job recruitment.
LLM alignment is fundamentally challenged by the dynamic and inconsistent nature of their internal "priority graphs," which adversaries can exploit through context manipulation.
Achieve up to 102% Sharpe Ratio improvement and 17.5% directional accuracy gain by unifying event-centric data construction and decision-oriented fine-tuning with a hierarchical gated reward model.
Naive application of LLM inference optimizations can *hurt* the performance of smaller reasoning models, highlighting the need for RLLM-specific serving strategies.