Search papers, labs, and topics across Lattice.
Affiliation:
5
0
6
21
Shifting from fixed-length tokens to semantic subwords can dramatically enhance the efficiency of attention mechanisms in generative recommenders.
SITA achieves target-aware compression that outperforms traditional methods, striking a balance between efficiency and specificity in long-sequence recommendations.
LLMs get a reasoning boost by treating information extraction not as a one-off task, but as a dynamic cache that persists and filters information across multiple steps.
Achieve personalized generation with cloud-scale reasoning while preserving user privacy, thanks to a novel asymmetric collaboration framework that's also 2x faster.
Forget complex model architectures for cross-domain recommendation: Taesar shows that cleverly transforming your data can unlock better performance with standard models.