Search papers, labs, and topics across Lattice.
Nanyang Technological University
2
0
4
LLM judges are swayed by style, mistaking it for substance, but a new framework can significantly mitigate this bias.
Even the best vision-language models struggle to reliably set fine-grained GUI states, achieving only 33% accuracy on a new benchmark, but targeted visual hints suggest a clear path to improvement.