Search papers, labs, and topics across Lattice.
1
0
2
LLMs can effectively assess voice-agent interactions, but their reliability hinges on specific metrics and evaluation setups, revealing a nuanced landscape for automated judgment.