Search papers, labs, and topics across Lattice.
Eastern Institute of Technology, Ningbo
6
0
9
7
Confidence in reasoning models can be dramatically improved by strategically leveraging position-aware signals, leading to better performance in challenging tasks.
AdaSR achieves superior reasoning performance in dynamic environments by enabling models to adaptively allocate computation during streaming input.
miniReranker slashes reranking runtime to under 1% of dense implementations while retaining over 96% of performance, revolutionizing efficiency in multimodal tasks.
Static depth pruning emerges as the most effective strategy for LLM acceleration, achieving near-theoretical speedup limits in memory-bounded contexts.
Low-KL agreement can trap models in ineffective training regimes, but KAT offers a dynamic solution that boosts accuracy while slashing rollout lengths.
FPGAs can beat GPUs at dynamically allocating computation for LLM inference, thanks to a new architecture that fuses operations, uses mixed precision, and caches KV values on-chip.