Search papers, labs, and topics across Lattice.
University of Wisconsin-Madison
5
0
11
4
GeoDES achieves unprecedented accuracy in storm structure synthesis, outperforming existing models and redefining weather prediction capabilities.
WARP reveals the hidden training data portfolios of foundation models with remarkable accuracy, challenging the opacity of model training processes.
Current AI agents only manage to complete 20.6% of complex real-world tasks, revealing a stark gap in their capabilities compared to human users.
Programmatic judges can match the performance of large LLMs while slashing costs and boosting evaluation speed by orders of magnitude.
Forget expensive human annotations: RubiCap uses LLM-generated rubrics to train image captioning models via RL, achieving superhuman performance and even improving VLM pretraining.