Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
Aggregation metrics in benchmarks can mask the irreplaceable strengths of models, leading to a misleading evaluation of their true capabilities.
Non-imperative syntactic structures can undermine safety alignment in large language models, exposing them to sophisticated jailbreaks.
Systematic feature engineering can outperform advanced model architectures, closing a critical evaluation gap in tabular benchmarks.