Research1 min read
PRAGMA: Benchmark for Personalized Guidance in Long Conversations
PRAGMA is a new benchmark designed to evaluate how well LLMs provide personalized guidance across extended conversations. Experiments show current systems struggle with evidence recovery and memory-grounded reasoning, indicating a need for improved conversational memory architectures.
From arXiv cs.AI