collaborators

5 papers

cs.CL2026

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling

Xing Yue, Linjuan Wu, Daoxin Zhang +2

Open-ended reward modeling requires judges that can follow subtle, domain-specific preferences when verifiable answers are unavailable. Existing rubric-based methods often address…

cs.CL2026

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

Linjuan Wu, Ruiqi Zhang, Xinze Lyu +7

Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its informal style, cultural referenc…

cs.LG2026

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control

Xincheng Yao, Ruoqi Li, Cheng Chen +4

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a pivotal technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, the de f…

cs.CL2026

Pause or Fabricate? Training Language Models for Grounded Reasoning

Yiwen Qiu, Linjuan Wu, Yizhou Liu +9

Large language models have achieved remarkable progress on complex reasoning tasks. However, they often implicitly fabricate information when inputs are incomplete, producing confi…

cs.CV2025

VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models

Jiacheng Ruan, Wenzhen Yuan, Xian Gao +6

Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Re…