LLMs1 min read
StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation
arXiv:2609.16340v1 Announce Type: new Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system's outputs, which the newer model may alr...
From arXiv cs.CL


