15 papers
PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment
Jiawei Feng, Jiancan Wu, Xingyu Zhu +3
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal se…
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
Kexin Huang, Junkang Wu, Jinda Lu +7
Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process remains notoriously fragile. In this wor…
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu +7
Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing RLVR method…
R^2-Mem: Reflective Experience for Memory Search
Xinyuan Wang, Wenyu Mao, Junkang Wu +2
Deep search has recently emerged as a promising paradigm for enabling agents to retrieve fine-grained historical information without heavy memory pre-managed. However, existing dee…
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
Jinda Lu, Junkang Wu, Jinghan Li +6
Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for multimodal large language models (MLLMs) have mainly focused on improving final answer correctness and…
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
Jinda Lu, Junkang Wu, Jinghan Li +6
Extending Reinforcement Learning with Verifiable Rewards (RLVR) to multimodal large language models (MLLMs) faces a fundamental challenge: their responses inherently interleave per…