Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
Yankai Yang, Yancheng Long, Hongyang Wei +12
Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of generative models. For complex tasks such as i…
cs.AI2026
When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel
Wenkai Li, Fan Yang, Ananya Hazarika +2
Chain-of-thought (CoT) traces are increasingly used both to improve language model capability and to audit model behavior, implicitly assuming that the visible trace remains synchr…
cs.AI2026
KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA
Fan Yang, Rui Meng, Trudi Di Qi +2
Reinforcement learning (RL) has emerged as a promising paradigm for inducing explicit reasoning behaviors in large language and vision-language models. However, reasoning-oriented…