3 citations · 7 across the 19 of their papers we have counts for
Showing cs.ROShow all
3 papers · 1 filter
cs.RO2026
AnchorVLA: Anchored Diffusion for Efficient End-to-End Mobile Manipulation
Jia Syuen Lim, Zhizhen Zhang, Peter Bohm +3
A central challenge in mobile manipulation is preserving multiple plausible action models while remaining reactive during execution. A bottle in a cluttered scene can often be appr…
cs.RO2025
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
Yuxia Fu, Zhizhen Zhang, Yuqi Zhang +3
Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a…
cs.RO2025
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
Zhizhen Zhang, Lei Zhu, Zhen Fang +2
Pre-training vision-language representations on human action videos has emerged as a promising approach to reduce reliance on large-scale expert demonstrations for training embodie…