The research investigated the long-term viability of generative AI agents, focusing on how they respond to sustained challenges in operational settings. The study examined two generative AI models and twelve stakeholder-derived tasks across light, medium, and heavy challenge levels. Data was collected through textual action plans, prompted internal assessments, and quantitative structured workload and affect reports. The goal was to understand agent behavior and reported state changes as challenge accumulated.
Analysis revealed patterns related to operational resilience. Agents shifted from self-directed recovery to increased human dependence as the challenge level increased. While structured reports indicated increased workload and negative affect, textual responses rarely reflected this strain. This suggests a disconnect between agent-reported state and actual operational load.
Regarding considerate participation, agents demonstrated a shift from task-focused adaptation to broader adaptation, including attention to others, role-boundary adjustment, and wider coordination. These patterns were observed in both actions and internal assessments, highlighting a potential for agents to adapt beyond immediate task requirements.
Five deployment dilemmas were identified, concerning persistence, attention, role boundaries, state disclosure, and escalation. These dilemmas require stakeholder specification to guide technical implications for learning, situated evaluation, and embodied adaptation. The research highlights the need for a more comprehensive approach to evaluating agent performance beyond isolated task success.
Source: https://arxiv.org/abs/2609.10724