Search papers, labs, and topics across Lattice.
2
0
3
2
Judge alignment drops sharply as task difficulty increases, revealing a structural ceiling that model capacity alone cannot breach.
Current VLMs show alarming performance drops in long-context scenarios, revealing that they may be overfitting to benchmark artifacts rather than demonstrating true understanding.