Search papers, labs, and topics across Lattice.
8
0
11
4
BPO achieves up to 6.1% higher success rates in sandbox-native RL tasks while cutting down on the number of required policy updates by 38%.
Traditional Top-K metrics fail to capture true user preferences, but the LLM-as-a-Judge framework offers a semantic approach that enhances both reliability and explainability in recommendation evaluations.
Explainable detection of hateful videos is now possible, revealing the nuanced reasoning behind classifications that traditional methods overlook.
Saturated LLM benchmarks can be revived without creating new datasets: a self-improving LLM judge in an elimination tournament recovers ranking signal and breaks ties.
Long-form video generation struggles with transitions, scoring only 0.256 on transition quality even when prompt fulfillment is high (0.71), revealing a critical bottleneck exposed by the new DirectorBench diagnostic benchmark.
Stop ignoring the future: adaptively weighting future user interactions during training can significantly boost sequential recommendation accuracy.
Forget retraining: ReAd dynamically adapts deployed sequential recommendation models to real-time preference shifts by retrieving and integrating collaborative user preference signals at test time.
Forget brute-force distillation: this method uses pedagogical principles to distill LLMs, boosting student model performance on complex reasoning tasks by up to 22.3%.