collaborators

6 papers

cs.AI2026

Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning

Wenxuan Zhang, Yuhui Wang, Donggang Jia +5

Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end t…

cs.LG2026

Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization

Yanning Dai, Yuhui Wang, Dylan R. Ashley +1

Morphology-control co-design concerns the coupled optimization of an agent's body structure and control policy. This problem exhibits a bi-level structure, where the control dynami…

cs.LG2026

A Unified Framework for Rethinking Policy Divergence Measures in GRPO

Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan +6

Reinforcement Learning with Verified Reward (RLVR) has emerged as a critical paradigm for advancing the reasoning capabilities of Large Language Models (LLMs). Most existing RLVR m…

cs.CV2025

RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding

Xiyan Liu, Han Wang, Yuhu Wang +4

Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous…

cs.LG2025

Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning

Yuhui Wang, Qingyuan Wu, Dylan R. Ashley +4

The Value Iteration Network (VIN) is an end-to-end differentiable neural network architecture for planning. It exhibits strong generalization to unseen domains by incorporating a d…

cs.LG2025

Directly Forecasting Belief for Reinforcement Learning with Delays

Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan +6

Reinforcement learning (RL) with delays is challenging as sensory perceptions lag behind the actual events: the RL agent needs to estimate the real state of its environment based o…