A study examined seven large language models, including frontier GPT-5.1, using the World Values Survey. The research found that incorporating demographic profiles improved value alignment accuracy for most models. However, this improvement was accompanied by a systematic reduction in the preservation of individual differences. This behavior is termed ‘alignment by stereotyping’. The study utilized permutation tests across 10,000 permutations and six demographic attributes to confirm this trend.
The analysis revealed that top-performing models compressed individuals significantly above the human baseline. Furthermore, within-family scaling amplified this tradeoff, negatively affecting intrinsic cultural understanding. The research utilized a synthetic dialogue dataset validated on real human-chatbot conversations from PRISM to further investigate this phenomenon.
Experiments showed that distributing demographic signals across conversational turns partially suppressed prototype retrieval compared to compact demographic labels, as validated through real conversations from PRISM. Replication at a larger scale is required to fully confirm these findings. The study highlights a potential tradeoff between personalization and individual expression in LLM deployments.
Source: https://arxiv.org/abs/2609.05993