Search papers, labs, and topics across Lattice.
1
0
2
3
Misleading benchmarks reveal that models can accept incorrect interpretations of event semantics, challenging the reliability of current NLI evaluations.