Search papers, labs, and topics across Lattice.
Beihang University
3
0
7
CrossPool achieves a staggering 10.4x reduction in tail latency for bursty long-context requests by decoupling model weights from KV-cache in GPU memory.
TabEmbed leapfrogs existing text embedding models to achieve SOTA performance on tabular data by reformulating tasks as semantic matching problems and using contrastive learning.
Ditch the manual feature engineering: KMLP's hybrid KAN-gMLP architecture automatically learns complex feature transformations and interactions, outperforming GBDTs on web-scale tabular data.