activity
20242026
collaborators

21 papers

cs.CL2026

User-Aware Active Knowledge Acquisition for Emotional Support Dialogue

Mufan Xu, Kehai Chen, Jiahao Hu +4

Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-turn interactions while leve…

cs.CL2026

Long-form RewardBench: Evaluating Reward Models for Long-form Generation

Hui Huang, Yancheng He, Wei Liu +7

The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models i…

cs.CL2026

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

Hongli Zhou, Hui Huang, Rui Zhang +5

Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by judgment biases. Accurately evalu…

cs.CV2026

Culture In a Frame: CB as a Comic-Based Benchmark for Multimodal Culturally Awareness

Yuchen Song, Andong Chen, Wenxin Zhu +4

Cultural awareness capabilities have emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their…

cs.CL2026

Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models

Mufan Xu, Kehai Chen, Xuefeng Bai +4

Existing policy-gradient methods for auto-regressive language models typically select subsequent tokens one at a time as actions in the policy. While effective for many generation…

cs.AI2026

Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling

Andong Chen, Wenxin Zhu, Qiuyu Ding +3

Chain-of-Thought reasoning has driven large language models to extend from thinking with text to thinking with images and videos. However, different modalities still have clear lim…