Search papers, labs, and topics across Lattice.
2
0
4
Just two factors can explain over 90% of a model's performance across 133 benchmarks, drastically simplifying evaluation processes.
Turns out, the terminal feedback your CLI agent throws away is actually a goldmine of dense supervision, allowing for significant performance gains and even self-improvement.