Research1 min read
MERIT Benchmark Evaluates Long-Term Memory in Tool-Using Agents
A new benchmark, MERIT, measures the cost-effectiveness of long-term memory for LLM agents executing tasks. Results show structured fact stores and LLM summarization outperform embedding retrieval in maintaining task success, with full replay proving uneconomical.
From arXiv cs.AI
