29 papers
STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models
Kasra Torshizi, Anukriti Singh, Sidharth Mathur +3
Vision-language-action (VLA) models have shown impressive generalization, but often lack interpretability and can struggle to follow precise natural language instructions that enco…
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
Rui Liu, Dian Yu, Tong Zheng +8
Reinforcement learning with verifiable rewards (RLVR) has advanced reasoning capabilities in multimodal large language models. However, existing methods typically treat visual inpu…
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
Rui Liu, Dian Yu, Zhenwen Liang +6
Aligning Multimodal Large Language Models (MLLMs) requires reliable reward models, yet existing single-step evaluators can suffer from lazy judging, exploiting language priors over…
Reinforcing Multimodal Reasoning Against Visual Degradation
Rui Liu, Dian Yu, Haolin Liu +6
Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies remain brittle against real-wor…
Zero Shot Coordination for Sparse Reward Tasks with Diverse Reward Shapings
Keenan Powell, Peihong Yu, Pratap Tokekar
Many Multi-Agent Reinforcement Learning (MARL) agents fail to adapt properly to cooperating with agents trained with the same objectives but different seeds, algorithms, or other t…
AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
Anukriti Singh, Kasra Torshizi, Khuzema Habib +3
Vision-based robot learning often relies on dense image or point-cloud inputs, which are computationally heavy and entangle irrelevant background features. Existing keypoint-based…