Search papers, labs, and topics across Lattice.
2
0
3
LLMs struggle with statistical analysis, achieving only 68.6% accuracy on a new benchmark designed to rigorously test their capabilities.
Advanced image editing models may look good but often miss the mark on logical consistency, revealing a critical gap in current AI capabilities.