Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
LLMs can achieve near-human performance on hard-labels but struggle with soft-labels, revealing a critical gap in automatic evaluation methods that NAPHA effectively bridges.
Task decomposition in LLMaJ doesn't boost NLG evaluation performance鈥攊t's the human labels that matter.