collaborators

15 papers

cs.CV2026

PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment

Jiawei Feng, Jiancan Wu, Xingyu Zhu +3

Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal se…

cs.LG2026

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

Kexin Huang, Junkang Wu, Jinda Lu +7

Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process remains notoriously fragile. In this wor…

cs.LG2026

Experience Augmented Policy Optimization for LLM Reasoning

Jinda Lu, Kexin Huang, Junkang Wu +7

Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing RLVR method…

cs.CL2026

R^2-Mem: Reflective Experience for Memory Search

Xinyuan Wang, Wenyu Mao, Junkang Wu +2

Deep search has recently emerged as a promising paradigm for enabling agents to retrieve fine-grained historical information without heavy memory pre-managed. However, existing dee…

cs.CV2026

Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR

Jinda Lu, Junkang Wu, Jinghan Li +6

Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for multimodal large language models (MLLMs) have mainly focused on improving final answer correctness and…

cs.CV2026

Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs

Jinda Lu, Junkang Wu, Jinghan Li +6

Extending Reinforcement Learning with Verifiable Rewards (RLVR) to multimodal large language models (MLLMs) faces a fundamental challenge: their responses inherently interleave per…