Search papers, labs, and topics across Lattice.
2
0
2
1
Even state-of-the-art models only achieve pass rates below 60% on a new benchmark that spans 1,431 diverse tasks, exposing critical weaknesses in general AI capabilities.
Stop wrestling with finicky evaluation codebases: One-Eval lets you specify LLM evaluation tasks in natural language and automatically executes them end-to-end.