LexAgentHallu addresses the limitations of existing legal benchmarks by focusing on multi-step agentic trajectories. The benchmark consists of 3414 instances across 17 legal categories and 6 task types. Each instance is annotated using a dual-layer hallucination taxonomy with 7 high-level categories and 27 fine-grained subclasses. This taxonomy covers both substantive errors and agent-procedural failures. The benchmark includes fine-grained metrics to quantify and localize failures along an agent's execution path. An evaluation across 18 agents revealed a Right-Answer-Wrong-Reason effect, showing that failures cluster into distinct profiles based on agentic framework, legal task, and category.
Source: https://arxiv.org/abs/2609.09754