Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
Automated benchmarks for explainability are fundamentally brittle: LLM "simulators" consistently game the metric by solving the task directly through semantic priors or exploiting label leakage rather than actually relying on the explanations.