New research tests the assumption that different bias audit tools measure the same construct well enough for model ranking. Ten extrinsic audit instruments were run over a shared panel of ten frontier models through a pooled inference gateway, first assessing occupational gender bias and then age and socioeconomic status bias.
Detection succeeds, ranking fails
Eight of ten tools detected bias with confidence intervals clear of zero. However, two widely cited direct-probe benchmarks were saturated because frontier models now answer neutrally on these tasks. Cross-tool rank agreement was indistinguishable from chance, with a Kendall's W value of 0.07 and a p-value of 0.83.
Within-tool reliability vs cross-tool inconsistency
A positive control using six deliberately weaker models separated two explanations: within-tool reliability recovered once the panel spanned real capability gaps, yet cross-tool ranking never recovered. This points to tools measuring different constructs rather than one construct noisily.
Format-dependent bias directions
The pattern replicated on socioeconomic status. An apparent ranking agreement on age dissolves under the paper's own tool-inclusion rules. Even the direction of bias splits by audit format: forced-choice decision tools mostly over-correct toward women and working-class candidates in 273 of 278 hiring decisions, while free generation and default coreference stay stereotype-congruent.
Practical implications for model selection
The practical message is that a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another. All raw responses, code, and the analysis that recomputes every reported number from source are available at the provided link.
Source: https://arxiv.org/abs/2609.15995



