activity
20242026
collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning

Rui Liu, Dian Yu, Tong Zheng +8

Reinforcement learning with verifiable rewards (RLVR) has advanced reasoning capabilities in multimodal large language models. However, existing methods typically treat visual inpu…

cs.AI2025

MMCD: Multi-Modal Collaborative Decision-Making for Connected Autonomy with Knowledge Distillation

Rui Liu, Zikang Wang, Peng Gao +3

Autonomous systems have advanced significantly, but challenges persist in accident-prone environments where robust decision-making is crucial. A single vehicle's limited sensor ran…

cs.AI2025

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach

Chak Lam Shek, Guangyao Shi, Pratap Tokekar

Multi-agent reinforcement learning (MARL) requires coordinated and stable policy updates among interacting agents. Heterogeneous-Agent Trust Region Policy Optimization (HATRPO) enf…

cs.AI2025

VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences

Anukriti Singh, Amisha Bhaskar, Peihong Yu +4

Designing reward functions for continuous-control robotics often leads to subtle misalignments or reward hacking, especially in complex tasks. Preference-based RL mitigates some of…

cs.AI2024

PLANRL: A Motion Planning and Imitation Learning Framework to Bootstrap Reinforcement Learning

Amisha Bhaskar, Zahiruddin Mahammad, Sachin R Jadhav +1

Reinforcement Learning (RL) has shown remarkable progress in simulation environments, yet its application to real-world robotic tasks remains limited due to challenges in explorati…