Search papers, labs, and topics across Lattice.
Affiliation:
2
0
2
2
Rubric-based grading allows small language models to outperform larger models, revealing that grading quality hinges more on structured criteria than on model size or judge intelligence.
ChartGenEval reveals that distinct evaluation signals can improve rhythm-game chart generation, challenging the reliance on traditional reconstruction metrics.