The ORQA framework was presented as a method for assessing the knowledge of large language models within professional contexts. The approach utilizes a pipeline combining automated processing with human review to generate question-answer pairs linked to O*NET occupations. These pairs are sourced from a diverse set of 187 websites, including regulatory agencies and professional organizations. The resulting question set covers 116 occupations across 21 major SOC groups, totaling 480 questions.
Testing was conducted on 15 state-of-the-art models, including Claude Opus 4.6, GPT-5.4 and Claude Sonnet 4.6. Performance varied significantly across occupations, with healthcare-related roles demonstrating the highest accuracy at approximately 78%, compared to 40% for Office and Administrative Support. Individual occupations, such as Sheet Metal Workers and Fish and Game Wardens, showed essentially zero performance.
Analysis of the results indicated that open-ended questions and weighting by wage bill did not substantially impact model rankings on the benchmark. The research suggests that leveraging existing, trusted occupation-specific information offers a scalable approach to evaluating LLM performance in professional domains. The data and code are available for public access.
The framework’s output includes a dashboard with interactive data visualization. The results and supporting data are available at the specified URL. Source: https://arxiv.org/abs/2609.12366



