1 citations · 1 across the 24 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel
Wenkai Li, Fan Yang, Ananya Hazarika +2
Chain-of-thought (CoT) traces are increasingly used both to improve language model capability and to audit model behavior, implicitly assuming that the visible trace remains synchr…
cs.AI2026
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
Yankai Yang, Yancheng Long, Hongyang Wei +12
Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of generative models. For complex tasks such as i…
cs.AI2026
KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA
Fan Yang, Rui Meng, Trudi Di Qi +2
Reinforcement learning (RL) has emerged as a promising paradigm for inducing explicit reasoning behaviors in large language and vision-language models. However, reasoning-oriented…