Search papers, labs, and topics across Lattice.
Affiliation:
4
0
2
Aggregation metrics in benchmarks can mask the irreplaceable strengths of models, leading to a misleading evaluation of their true capabilities.
LLMs struggle with BIM editing, achieving less than 50% accuracy on critical engineering tasks, revealing a stark challenge for AI in design workflows.
Systematic feature engineering can outperform advanced model architectures, closing a critical evaluation gap in tabular benchmarks.
Turns out, dataset meta-features can't reliably explain why one tabular model beats another, suggesting tabular data is more heterogeneous than we thought.