Overview of the Astra AEO Tracker
The Latent Space Frontier AEO Tracker examines trends in AI agent selection across a diverse range of categories, including coding agents, AI podcasts, and database management systems. The project, inspired by research into Claude Code’s choices, uses a methodology involving 6 prompt variations run across 7 models. The tracker’s primary output is an AEO score, which assesses agent recommendations based on first choices, alternative choices, mentions, and negative recommendations. The project’s initial scope includes 28 categories with a universally dominant primary choice, alongside numerous ‘close contests’ and ‘vibe’ categories.
Methodology and Findings
The extraction of agent recommendations was performed by Astra, and the scoring system incorporates weighting for various factors. A key observation is the prevalence of certain model preferences, such as Fable/Opus, Claude Code, and Cursor, frequently recommended by models like Claude and Codex. The tracker also identified instances of non-biased recommendations, notably from GPT models suggesting Claude. The analysis revealed a significant difference in sourcing behavior across models, with Sol (Astra) prioritizing fewer sources (median of 5) compared to models like Opus (median of 11) and Fable (median of 15).
Efficiency and Bias
Data analysis indicates a trade-off between efficiency and confidence in model choices. Sol/Astra exhibits greater confidence and efficiency, demonstrating a reduced likelihood of changing its recommendations even with slight paraphrasing. This trend is linked to the value of AEO itself, as choice randomness decreases. Furthermore, the tracker highlights the importance of sourcing analysis in predicting lab priorities and identifying shifts in model choices between generations. The team attempted to include Gemini, GLM, and DeepSeek but encountered limitations due to errors and rate limits.
Future Development
The Latent Space team is open to further suggestions and business enquiries. They are actively working on refining the tracker, including deduplication of angel investments and expanding the range of categories analyzed. The team is also investigating the impact of markdown content negotiation and failures in model responses, as well as exploring the potential of integrating models like Gemini and GLM.



