Research introduced environment-probing curation, an extension for asynchronous curator agents. This extension provides least-privilege, read-only world tools for checking, scoping, and refreshing candidate memories. The approach requires no model retraining and maintains the existing task agent, retriever, memory representation, and production write authority. Experiments were conducted using a GitHub Copilot (GHCP) harness on the CLBench database and 90 adapted APEX management-consulting tasks. Results showed a pass rate increase from 39% to 73% on CLBench, alongside reduced queries and task-agent costs. Across six APEX worlds, all memory-versus-baseline mean reward comparisons were positive, with probing delivering the best reward gain per dollar in five worlds. Furthermore, probing achieved higher mean reward than GHCP + Mem on Sonnet 4.6 and Opus 4.7 without schema drift. This approach transforms agent memory curation into an environment-informed, auditable process.
Source: https://arxiv.org/abs/2609.11060