The research introduces RESCUE, a benchmark constructed from real couple and family interview conversations. It contains 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video data. The benchmark defines six tasks focused on ‘Relation-Aware Emotional Support Conversation’ evaluating ‘Relational Understanding’ and ‘Relation-Sensitive Support’. Ten LLMs were tested on these tasks. Results indicate that current models perform adequately when relying on immediate emotional or intervention cues. However, performance degrades significantly on tasks requiring the prediction of relation patterns, viewpoints, and support strategies. This suggests a current inability of LLMs to effectively model interpersonal dynamics. The findings highlight the need for improvements in LLM architectures to better handle the complexities of multi-party conversations and relation-sensitive support.
Source: https://arxiv.org/abs/2609.09657