Search papers, labs, and topics across Lattice.
This study critically evaluates the effectiveness of task decomposition in the LLMaJ framework for automatic NLG evaluation by comparing performance across multiple datasets. Contrary to prior assumptions, the findings reveal that task decomposition does not enhance performance over a baseline method that avoids decomposition. Instead, the observed improvements in previous studies were attributed to the use of human labels for training rather than the decomposition approach itself, suggesting that LLMaJ can match human annotators' performance without this added complexity.
Task decomposition in LLMaJ doesn't boost NLG evaluation performance鈥攊t's the human labels that matter.
The LLM-as-a-judge (LLMaJ) framework has emerged as a promising solution for cheap, reproducible, reference-free Natural Language Generation (NLG) evaluation. Prior work seeks to improve LLMaJ by decomposing evaluation tasks into simpler sub-tasks. In this work, we systematically compare LLMaJ methods with and without decomposition on multiple NLG datasets. We find no evidence that LLMaJ with task decomposition leads to performance gains over a fair baseline that does not use decomposition. Instead, we find that previously reported performance gains in decomposition-based LLMaJ stem from using human labels as training data, and not task decomposition itself. Also, we find that, when human labels are available, LLMaJ without using task decomposition can perform comparably to human annotators.