Search papers, labs, and topics across Lattice.
4
0
10
6
WARP reveals the hidden training data portfolios of foundation models with remarkable accuracy, challenging the opacity of model training processes.
Programmatic judges can match the performance of large LLMs while slashing costs and boosting evaluation speed by orders of magnitude.
Optimal LLM pretraining actually requires *overtraining* when you account for inference costs, overturning conventional scaling wisdom.
Forget expensive human annotations: RubiCap uses LLM-generated rubrics to train image captioning models via RL, achieving superhuman performance and even improving VLM pretraining.