Search papers, labs, and topics across Lattice.
GEB-Bench is a novel benchmark designed to evaluate models' ability to recognize and map abstract structural motifs across different representations, such as natural scenes, folk stories, mathematical theorems, and programmatic skeletons. The study reveals a significant gap in models' performance, where they excel at recognizing structures within a single voice but struggle to transfer that understanding across different voices, indicating a lawful abstraction failure. Key findings show that errors are more closely related to the formal geometry of the motifs than to perceptual geometries, with frontier models consistently converging on incorrect answers despite increased capacity.
Models can identify abstract structures in isolation but fail dramatically when asked to translate that understanding across different contexts, revealing a critical gap in AI's reasoning abilities.
Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-reference, a strange loop, a Mobius twist--in the spirit of Godel, Escher, Bach. Each motif is told in several voices: a natural scene whose composition is the structure, a folk story whose telling enacts it through a mechanically checkable form device, a mathematical theorem, and a programmatic skeleton; surface parameters are declared nuisance variables and never scored. Motifs, voices, and the structural changes between them form a small cross-modal category, and GEB-Bench's tasks are its questions. Evaluating twelve open and proprietary models, we find that abstraction failure is lawful. The central finding is a gap between recognition and cross-voice mapping: models identify a structure within one voice far better than they carry it across voices; every model pays this tax, and mapping strong enough to narrow it appears only at the frontier tier. Two patterns support it. Errors align more strongly with the designed formal geometry than with measured perceptual geometries, and frontier models from different vendors converge on the same wrong answers; and surface complexity taxes every model that reads structure, with capacity buying headroom rather than immunity. GEB-Bench is fully generative and released with its pipeline.