collaborators

8 papers

cs.LG2026

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Jing Liang, Hongyao Tang, Yi Ma +9

Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. O…

cs.RO2026

From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation

Yifu Yuan, Haiqin Cui, Yibin Chen +7

Achieving generalization in robotic manipulation remains a critical challenge, particularly for unseen scenarios and novel tasks. Current Vision-Language-Action (VLA) models, while…

cs.AI2025

Key Decision-Makers in Multi-Agent Debates: Who Holds the Power?

Qian Zhang, Yan Zheng, Jinyi Liu +2

Recent studies on LLM agent scaling have highlighted the potential of Multi-Agent Debate (MAD) to enhance reasoning abilities. However, the critical aspect of role allocation strat…

cs.LG2025

DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question Answering

Rong Cheng, Jinyi Liu, Yan Zheng +6

Multi-Hop Question Answering (MHQA) tasks permeate real-world applications, posing challenges in orchestrating multi-step reasoning across diverse knowledge domains. While existing…

cs.LG2025

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model

Jing Liang, Hongyao Tang, Yi Ma +5

Reinforcement Learning (RL) has demonstrated its potential to improve the reasoning ability of Large Language Models (LLMs). One major limitation of most existing Reinforcement Fin…

cs.LG2025

Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning

Jinyi Liu, Yi Ma, Jianye Hao +4

In recent years, data-driven reinforcement learning (RL), also known as offline RL, have gained significant attention. However, the role of data sampling techniques in offline RL h…