Search papers, labs, and topics across Lattice.
9
3
13
5
A unified framework reveals that the choice of tokenization and vocabulary topology can significantly influence the performance of discrete diffusion models, unlocking new avenues for optimization.
BPO achieves up to 6.1% higher success rates in sandbox-native RL tasks while cutting down on the number of required policy updates by 38%.
Span-level hallucination detection can now effectively address structured inputs like code and tool outputs, not just natural language, revealing a critical gap in current RAG evaluations.
Set representation models can be made robust to inference-time corruptions like outliers and missing data by training against a learned barycentric adversary.
Even state-of-the-art vision-language models frequently lie and hallucinate when playing social deduction games, raising serious questions about their reliability in real-world applications requiring grounded reasoning.
Stop overpaying for LLM serving: intelligently routing requests to specialized pools based on token budget slashes GPU costs by up to 42% and dramatically improves reliability.
DANCEMATCH enables efficient large-scale dance retrieval by creating compact, discrete motion signatures that capture the spatio-temporal structure of dance, moving beyond continuous embeddings.
LLM GPU fleets can be analytically optimized into a two-pool architecture with gateway-layer compression, slashing costs by up to 82% without sacrificing latency.
Seemingly idle LLM inference fleets can be secretly broken, and this simulator helps you find out why before you buy.