Search papers, labs, and topics across Lattice.
4
0
4
8
This work defines the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference, and proposes Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales.
Simulating 8.3 billion diverse personas reveals nuanced user interactions that traditional evaluations miss, transforming how we assess AI systems.
Retaining the right contextual information can boost long-context modeling performance by over 20% compared to traditional methods.
RoboDojo reveals that integrating simulation and real-world tasks can significantly enhance the evaluation of robot manipulation policies, bridging the gap between theoretical performance and practical deployment.