The work introduces an offline self-generated framework for narrow-boundary safety in language models, combining controlled topic generation, coverage repair, and harmful-benign pairs for training and evaluation.
Training on refusal data through an escalation process significantly increases the model's ability to refuse targeted political persuasion, raising refusal rates from around 9% to over 84%, and reduces unsafe responses across benchmarks.
Results indicate that data composition influences the safety and usability trade-off, with boundary-pair data reducing over-refusal on held-out pairs, while harmful-side refusal remains high. Evaluating safety should consider both sides of the refusal boundary.
This approach offers insights into controlling safety boundaries in models, balancing safety and usability in deployment.
Source: https://arxiv.org/abs/2609.04482