Search papers, labs, and topics across Lattice.
2
0
3
0
Machine translation benchmarks have functionally saturated, but pairing human-authored failure cases with deterministic verification rules reveals critical multimodal blind spots that automated metrics consistently miss.
LLMs can mimic CBT dialogues, but still fail to convey the empathy and consistency of a human therapist.