LLMs1 min read
Benchmark for Social Pragmatic Inference in Chinese Online Comments
A new benchmark evaluates whether large language models can interpret social meaning in Chinese online comments, focusing on indirect and playful language. The strongest model achieves 81.42% accuracy.
From arXiv cs.CL