PRAGMA is a benchmark introduced to assess personalized guidance in long-term conversations involving large language models. The benchmark includes curated longitudinal conversation histories, evidence annotations, and guidance scenarios. These scenarios are grounded in evolving user contexts and address incorrect user assumptions. The benchmark focuses on evaluating retrieval systems, memory systems, and long-context models. Experiments across these systems reveal difficulties in recovering appropriate conversational evidence and effectively utilizing it for personalized guidance.
Source: https://arxiv.org/abs/2609.09664