4 papers
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
Yuhang Han, Yuyang Wu, Zhengbo Jiao +6
Reinforcement Learning from Verifiable Rewards (RLVR) has substantially enhanced the reasoning capabilities of large language models in abstract reasoning tasks. However, its appli…
OpenMic: A Multi-Agent-Based Stand-Up Comedy Generation System
Yuyang Wu, Hanzhong Cao, Jianhao Chen +1
Chinese stand-up comedy generation goes beyond plain text generation, requiring culturally grounded humor, precise timing, stage-performance cues, and implicit multi-step reasoning…
Momentum Posterior Regularization for Multi-hop Dense Retrieval
Zehua Xia, Yuyang Wu, Yiyun Xia +1
Multi-hop question answering (QA) often requires sequential retrieval (multi-hop retrieval), where each hop retrieves missing knowledge based on information from previous hops. To…
When More is Less: Understanding Chain-of-Thought Length in LLMs
Yuyang Wu, Yifei Wang, Ziyu Ye +3
Large Language Models (LLMs) employ Chain-of-Thought (CoT) reasoning to deconstruct complex problems. While longer CoTs are often presumed superior, this paper challenges that noti…