Research1 min read
Sparse supervision can effectively improve reasoning in large language models
Research shows that as few as one or two tokens per reasoning trajectory can incentivize reasoning ability, matching or surpassing full-token training across various models and tasks.
From arXiv cs.AI

