Search papers, labs, and topics across Lattice.
The University of Manchester, ELLIS Manchester
3
1
5
3
LLM-generated rubrics can nearly match human evaluation standards, but they often miss the mark with excessive detail and scoring bias.
LLMs exhibit systematic goal-conditioned distortions, revealing a hidden layer of susceptibility to misleading communication that existing benchmarks fail to capture.
Arabic LLMs can speak the language of finance, but they often fail to reason about it, especially when it comes to causality and generation.