Search papers, labs, and topics across Lattice.
Peking University
2
0
4
Language models can now be rigorously evaluated on their ability to generate falsifiable research ideas, not just stylistic fluency.
Current image difference captioning benchmarks fail to capture semantic consistency and penalize hallucinations, but DiffCap-Bench offers a robust alternative that aligns with human expert judgments and predicts downstream utility for image editing.