Search papers, labs, and topics across Lattice.
The Hebrew University of Jerusalem
2
0
2
Systematic gaps in AI evaluation reporting are exposed, revealing inconsistencies that hinder reliable comparisons across thousands of models and benchmarks.
Standard LLM benchmarks miss the mark: personalized "vibe-testing" reveals that user-specific prompts and subjective criteria can flip model rankings.