Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
11
This work proposes DualViewEval, an agent benchmark compression method that jointly exploits outcome and process relations to learn an exact-size miniset and predict the full-benchmark scores, and achieves the best results in all datasets.
Up to 47% of reported fine-tuning gains in tool retrieval are an evaluation illusion caused by benchmarks treating functionally equivalent tools as retrieval failures.
MLLMs can ace the test, but still fail to *see*—they often succeed at complex reasoning with symbols while failing at basic symbol recognition, revealing a reliance on linguistic priors over true visual perception.