activity
20242026
collaborators

10 papers

cs.CL2026

Adaptive Latent Agentic Reasoning

Dongwon Jung, Peng Shi, Yi Zhang +2

Large reasoning models improve performance by generating extended chain-of-thought (CoT) reasoning, but this behavior becomes inefficient when applied to LLM agents. Current LLM ag…

cs.CL2026

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

Qiyao Ma, Dechen Gao, Rui Cai +4

Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing d…

cs.CV2026

VITA: Vision-to-Action Flow Matching Policy

Dechen Gao, Boqi Zhao, Andrew Lee +6

Conventional flow matching and diffusion-based policies sample via iterative denoising from standard noise distributions (e.g., Gaussian), and require conditioning modules to repea…

cs.AI2025

GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective

Hang Wang, Junshan Zhang

Multi-agent reinforcement learning faces fundamental challenges that conventional approaches have failed to overcome: exponentially growing joint action spaces, non-stationary envi…

cs.RO2025

Ego-centric Learning of Communicative World Models for Autonomous Driving

Hang Wang, Dechen Gao, Junshan Zhang

We study multi-agent reinforcement learning (MARL) for tasks in complex high-dimensional environments, such as autonomous driving. MARL is known to suffer from the \textit{partial…

cs.RO2025

IN-RIL: Interleaved Reinforcement and Imitation Learning for Policy Fine-Tuning

Dechen Gao, Hang Wang, Hanchu Zhou +5

Imitation learning (IL) and reinforcement learning (RL) each offer distinct advantages for robotics policy learning: IL provides stable learning from demonstrations, and RL promote…