Search papers, labs, and topics across Lattice.
As frontier machine translation systems saturate standard benchmarks and game automated metrics, the authors construct the Last Translation Benchmark (LTB), a dynamic multimodal testbed designed explicitly to expose edge-case failure modes. Rather than relying on aggregate scoring or subjective human audits, LTB pairs human-authored adversarial inputs across text, audio, image, and video with handcrafted, programmatic verification rules targeting concrete errors. The resulting living framework provides an actionable, continuously updated diagnostic suite that tracks true model capabilities as translation systems evolve.
Machine translation benchmarks have functionally saturated, but pairing human-authored failure cases with deterministic verification rules reveals critical multimodal blind spots that automated metrics consistently miss.
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.