Search papers, labs, and topics across Lattice.
Case Western Reserve University
3
0
6
Test-time scaling can significantly enhance LLM reasoning capabilities, but without clear protocols, results are often incomparable and misleading.
Hybrid-thinking LLMs can be dramatically improved by simply separating the feed-forward pathways for reasoning and non-reasoning modes, leading to less leakage and better accuracy.
Agent evaluation is bottlenecked by environment interaction overhead, but ACE-Bench slashes this by using static JSON files, enabling fast and reproducible training-time validation.