Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
Rui Liu, Dian Yu, Tong Zheng +8
Reinforcement learning with verifiable rewards (RLVR) has advanced reasoning capabilities in multimodal large language models. However, existing methods typically treat visual inpu…
cs.AI2025★ 1 cited
Scaling Reinforcement Learning for Content Moderation with Large Language Models
Hamed Firooz, Rui Liu, Yuchen Lu +15
Content moderation at scale remains one of the most pressing challenges in today's digital ecosystem, where billions of user- and AI-generated artifacts must be continuously evalua…
cs.AI2025
MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
Lanxue Zhang, Yuqiang Xie, Fang Fang +3
Large Language Models demonstrate strong reasoning capabilities, which can be effectively compressed into smaller models. However, existing datasets and fine-tuning approaches stil…