Large language models (LLMs) are increasingly deployed in real-world systems such as financial markets. Improving individual model capabilities can sometimes degrade system-level outcomes due to shared training and architectures.
The study develops a framework showing how correlated actions among capable LLMs create a non-diversifiable risk floor. An agent-based simulation tested these predictions with LLM traders of varying capabilities.
Results indicate that more capable LLMs exhibit significantly correlated behaviors that grow with their capability. When their reasoning is accurate, increased participation reduces market risk, but shared misinformation environments turn this correlation into a liability.
These findings reveal a capability paradox: enhancing individual models does not necessarily improve overall system stability. It remains an open question whether similar dynamics occur in other domains.
Source: https://arxiv.org/abs/2609.04373