A benchmark was introduced to assess if large language models can recover pragmatic meanings in Chinese online comments that are indirect or playful. The dataset includes over 4,700 diagnostic items derived from more than 200,000 social media records, pairing comments with context and potential misreadings.
Eight models were evaluated as question writers and solvers in a cross-writer setting. The task proved challenging, with the best model reaching 81.42% accuracy when the writer was not the same as the solver. Overall, models achieved a mean accuracy of 68.70%, compared to human accuracy of 90.8%.
Results indicate that models can recognize broad irony or playfulness but often misidentify specific interactional mechanisms. This benchmark highlights the difficulty of interpreting nuanced social language in Chinese online communication, relevant for deploying models in social media analysis or conversational agents.
Source: https://arxiv.org/abs/2609.04384