activity
20242026
collaborators

31 papers

cs.CL2026

From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning

Yudong Wang, Zhe Yang, Wenhan Ma +6

Rewards that penalize unsupported claims can improve grounding in long-form generation, but they can also teach models to answer less. We study this refusal-to-richness trade-off i…

cs.CL2026

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Wenhan Ma, Jianyu Wei, Liang Zhao +10

Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains…

cs.CR2026

HauntAttack: When Attack Follows Reasoning as a Shadow

Jingyuan Ma, Rui Li, Zheng Li +4

Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However, the enhancement of reasoning abilities a…

cs.CL2026

SenseJudge: Human-Centric Preference-Driven Judgment Framework

Rui Li, Junfeng Liu, Xiangwen Kong +2

Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However, existing judgment approach…

cs.AI2026

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Bowen Ye, Rang Li, Qibin Yang +10

Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by…

cs.AI2026

SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning

Zheng Li, Qingxiu Dong, Jingyuan Ma +3

Recently, large reasoning models demonstrate exceptional performance on various tasks. However, reasoning models always consume excessive tokens even for simple queries, leading to…