LLMs1 min read
Document-Level MT Evaluation Shows Statistical Equivalence
Research found that document-level machine translation evaluation, presenting full documents to annotators, yields statistically equivalent scores and rankings compared to segment-level evaluations. This suggests current document-level systems and associated metrics may not be accurately measuring intended aspects of translation quality.
From arXiv cs.CL