Search papers, labs, and topics across Lattice.
3
0
5
9
Evolving rubrics from a single query can dramatically enhance LLM evaluation by eliminating reliance on external annotations and improving answer quality discrimination.
Skip the costly full training runs: this new metric accurately predicts face recognition dataset quality using only lightweight proxy models.
Observational user feedback, often dismissed as too noisy and biased, can actually power effective RLHF with the right causal modeling, achieving a 49.2% gain on WildGuardMix.