A new distillation method, correctness-gated multi-teacher distillation, was evaluated using an 63.9M-parameter student model. The student was trained on 4,330 sources and utilized a fixed three-response pool from seven teacher-based arms. Results indicated a relative accuracy increase of +0.1660 (95% observed-matrix interval [0.0670, 0.2455]) and a five-label macro-F1 increase of +0.1323 ([0.0916, 0.1731]) compared to unfiltered distillation. However, the weighted arm exhibited a decrease in the task-defined conditional unsafe-action rate by -0.4979 ([-0.5926, -0.3686]).
During an availability-amended audit at one reference seed, the weighted arm had zero Refuted recall and two seeds assigned NotEnoughInfo to all 167 claim examples. The audit compared weighted and unfiltered outputs, revealing a discrepancy in evidence-supported positives and unsupported material. The audit was conducted with non-paired samples, without serialized source overlap, and followed automatic summarization before annotation.
Hard filtering achieved 0.660 accuracy, 0.530 macro-F1, and 0.135 conditional unsafe rate. The implemented weighted arm showed no demonstrated incremental decision benefit over hard filtering. The experiment highlights decision redistribution with lost label functionality.
Source: https://arxiv.org/abs/2609.09702