Search papers, labs, and topics across Lattice.
3
0
4
0
Machine translation benchmarks have functionally saturated, but pairing human-authored failure cases with deterministic verification rules reveals critical multimodal blind spots that automated metrics consistently miss.
Optimal granularity in RAG benchmarks varies by dimension, with question complexity thriving on fine distinctions while other factors favor medium granularity.
You can accurately predict the NDCG of a 1B-parameter reranking model by only training models up to 400M parameters, unlocking massive compute savings.