RobustSGPO addresses limitations in semantic-gradient-based prompt optimization (SGPO) by controlling the edit scope and operation within agent harnesses. The system specifies the requested edit, constructs and checks the patch, and continues the search process from either the incumbent or retained snapshots. Evaluation utilized 120 tasks, 95 runs, and 7,350 candidate attempts within the AgentX brainstorming workflow. Periodic $1 o2 o3$ scheduling demonstrated a 0.28 test-score point improvement over fixed maximum permission. RobustSGPO increased completion on 30 held-out tasks from 60.0% to 80.0% and improved test quality from 3.77 to 4.14 with a 20-million-token budget. Task-family transfer and cumulative controls were also investigated. Category retention reduced source-task degradation after a shift, while random retention reached a higher destination endpoint. This approach offers measurable retention overhead through search-space control.
Source: https://arxiv.org/abs/2609.09646