Search papers, labs, and topics across Lattice.
Affiliation:
3
0
7
Foundation models already encode whether they are right or wrong: a lightweight attention layer over frozen internal token states reliably predicts classification errors across both text and multimodal domains.
Achieving up to 46% token compression without sacrificing accuracy, HMPO revolutionizes the efficiency of chain-of-thought reasoning in large language models.
Unified multimodal models suffer from internal conflict, but this work shows how to turn that interference into a surprisingly effective source of performance gains.