Search papers, labs, and topics across Lattice.
3
0
5
3
RubricRM adapts evaluation criteria dynamically, leading to significant performance gains in visual generative tasks compared to static reward models.
Even top-performing multimodal judges fail to reliably detect errors in specific modalities, revealing hidden blind spots in their evaluations.
Failure-driven post-training, combined with a meticulously curated 10M token STEM dataset, unlocks a 4.68% performance boost in LLM reasoning, proving that strategic data synthesis around model weaknesses is a powerful path to improvement.