Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
3
High reasoning accuracy in VLMs doesn't equate to reliability, as shown by GPT-5.2's 96% hallucination rate despite top performance metrics.
Annotation reporting in NLP is improving, but critical details that affect validity are still frequently overlooked, especially in model evaluations.
A clever RL setup lets small, open-source models beat GPT-4o at generating scientific figures from text, proving that data quality and reward engineering can trump sheer size.