Search papers, labs, and topics across Lattice.
3
0
6
Simulating 8.3 billion diverse personas reveals nuanced user interactions that traditional evaluations miss, transforming how we assess AI systems.
Over a quarter of tasks in popular AI benchmarks contain critical flaws that distort model evaluations, and this automated auditing framework can catch them.
Language models can bootstrap their reasoning abilities without human labels by learning from each other's aggregated answers, achieving significant gains in mathematical reasoning.