The paper introduces CORE, a method designed to address persona drift in large language models. CORE separates turn-local evidence from persistent persona-state revision, selectively updating grounded user preferences based on uncertainty-aware belief revision. The system aims to prevent updates driven by transient or ambiguous observations.
The research introduces PERSIST, a held-out benchmark specifically created to assess the robustness of persona-state revisions during prolonged, multi-turn interactions. This benchmark covers scenarios including ambiguity, conflict, and controlled social influence. The benchmark utilizes ALOE, PersonaChat, and PERSIST to evaluate CORE’s performance.
Across these datasets, CORE demonstrated improvements in personalized alignment and robustness, as measured by normalized closed-slot state fidelity. Human evaluation and mechanistic controls confirmed the system’s ability to explicitly control update decisions, surpassing improvements achieved through stronger generation or persistent memory alone.
The CORE system’s effectiveness was evaluated across three models: ALOE, PersonaChat, and the PERSIST benchmark. The research suggests that CORE improves personalized alignment and robustness, with complementary gains in normalized closed-slot state fidelity. Source: https://arxiv.org/abs/2609.12373



