Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model
Md Selim Sarowar, Md Tanvir Islam, Sungho Kim +1
Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure, and augmenting them with dense monocular depth i…
cs.RO2026
SUREFlow: State-space Uncertainty-aware REsidual Flow Matching for Robust Robot Manipulation
Md Tanvir Islam, Sai Navaneet Peddapalli, Sangmoon Lee +1
Generative vision-language-action policies have advanced robot manipulation, but they often exhibit instability under noise, partial observability, and stochastic initial condition…