3 papers
cs.RO2026
GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model
Md Selim Sarowar, Md Tanvir Islam, Sungho Kim +1
Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure, and augmenting them with dense monocular depth i…
cs.CV2026
Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval
Kyeongmo Chae, Jihoon Lee, Sangtae Ahn
This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task rather than traditional dete…
cs.RO2026
SUREFlow: State-space Uncertainty-aware REsidual Flow Matching for Robust Robot Manipulation
Md Tanvir Islam, Sai Navaneet Peddapalli, Sangmoon Lee +1
Generative vision-language-action policies have advanced robot manipulation, but they often exhibit instability under noise, partial observability, and stochastic initial condition…