4 papers
ROI-Driven Foveated Attention for Unified Egocentric Representations in Vision-Language-Action Systems
Xinhai Sun, Xiang Shi, Menglin Zou +1
The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action…
SaiVLA-0: Cerebrum--Pons--Cerebellum Tripartite Architecture for Compute-Aware Vision-Language-Action
Xiang Shi, Wenlong Huang, Menglin Zou +1
We revisit Vision-Language-Action through a neuroscience-inspired triad. Biologically, the Cerebrum provides stable high-level multimodal priors and remains frozen; the Pons Adapte…
Geometry-Aware Sparse Depth Sampling for High-Fidelity RGB-D Depth Completion in Robotic Systems
Tony Salloom, Dandi Zhou, Xinhai Sun
Accurate three-dimensional perception is essential for modern industrial robotic systems that perform manipulation, inspection, and navigation tasks. RGB-D and stereo vision sensor…
MT-Depth: Multi-task Instance feature analysis for the Depth Completion
Abdul Haseeb Nizamani, Dandi Zhou, Xinhai Sun
Depth completion plays a vital role in 3D perception systems, especially in scenarios where sparse depth data must be densified for tasks such as autonomous driving, robotics, and…