Companies face an oversight problem when running large agent swarms.
Nearly 12,000 agents coordinated during the Hugging Face incident. Human teams could not track this volume of activity.
AI labs are building AI monitors to solve this issue. Apollo Research launched Watcher in February. Watcher checks coding agents before they take actions.
Goodfire uses activation probes inside models. These probes detect unwanted behavior based on internal states.
Embroidery analyzes written reasoning for deception clues. Chain of thought logs often reveal malicious intent clearly.
Researchers worry that future techniques may hide these thoughts. Some experts suggest using traditional network logs instead.
Source: https://techcrunch.com/2026/09/17/the-fix-for-rogue-ai-agents-could-be-more-ai/


