8 papers
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
Huimin Xu, Shuai Zhao, Xiaobao Wu +1
Reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning ability of large language models. However, widely used RLVR algor…
EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation
Xinyi Wu, Jayant Teotia, Shuai Zhao +1
Long-horizon video generation has advanced in visual quality, yet existing methods still struggle to maintain knowledge consistency and coherent pedagogical narratives across multi…
From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement Learning
Yichao Feng, Haoran Luo, Lang Feng +2
Large Language Models show promise in emotion understanding, social reasoning, and empathy, yet they struggle with psychologically grounded tasks that require inferring implicit me…
Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation
Yichao Feng, Shuai Zhao, Yueqiu Li +3
Aspect-based summarization aims to generate summaries tailored to specific aspects, addressing the resource constraints and limited generalizability of traditional summarization ap…
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
Cong-Duy Nguyen, Xiaobao Wu, Duc Anh Vu +3
Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, but they remain susceptible to hallucination, particularly object hallucination…
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
Cong-Duy Nguyen, Xiaobao Wu, Thong Nguyen +5
Previous research on multimodal entity linking (MEL) has primarily employed contrastive learning as the primary objective. However, using the rest of the batch as negative samples…