Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
8
Task-CoEvolve slashes evaluation costs by 80% while maintaining performance parity with full-set validation in LLM harness optimization.
Even the most advanced AI models struggle with urban navigation, achieving only 17.1% accuracy compared to human performance of 77.3%.
AI-written papers are surprisingly prone to hallucination, with even state-of-the-art models like ClaudeCode averaging over 10 factual errors per paper despite strong presentation.