Standard translation benchmarks are saturating and automatic metrics are unreliable—this benchmark uses human-authored hard cases with explicit failure rules to provide reproducible evaluation that actually identifies what models get wrong.
This paper introduces the Last Translation Benchmark, a curated collection of challenging translation examples (text, images, audio, video) designed to expose weaknesses in state-of-the-art machine translation models.