Search papers, labs, and topics across Lattice.
2
0
3
This work analyzes over 790,000 translation outputs from 12 LLMs across 22 language pairs (LPs) and identifies 12 recurring noise patterns, which are group into formatting and content noise, and constructs TransClean, a controlled benchmark of 9,900 pairs of noisy and clean translation outputs.
Machine translation benchmarks have functionally saturated, but pairing human-authored failure cases with deterministic verification rules reveals critical multimodal blind spots that automated metrics consistently miss.