Search papers, labs, and topics across Lattice.
4
0
6
1
Evolving rubrics from a single query can dramatically enhance LLM evaluation by eliminating reliance on external annotations and improving answer quality discrimination.
Achieving 71.6% accuracy in diagnosing unseen medical conditions with just two labeled examples showcases the power of federated learning in low-resource clinical settings.
SafeRun achieves perfect safety in LLM-based running planning, a critical advancement for applications where safety violations can have serious consequences.
Observational user feedback, often dismissed as too noisy and biased, can actually power effective RLHF with the right causal modeling, achieving a 49.2% gain on WildGuardMix.