From the 1 of 5 linked papers with an AI index.
5 papers
VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models
Ruiqi Xian, Yuehan Xian, Jing Liang +2
The paper introduces VISA, a training-time method that uses a visual‑language model to audit and correct semantic labels of 3D voxel occupancy maps, improving object and rare‑class…
SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model
Zewei Zhou, Ruining Yang, Xuewei +8
Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. Howe…
GaussianSSC: Triplane-Guided Directional Gaussian Fields for 3D Semantic Completion
Ruiqi Xian, Jing Liang, He Yin +2
We present \emph{GaussianSSC}, a two-stage, grid-native and triplane-guided approach to semantic scene completion (SSC) that injects the benefits of Gaussians without replacing the…
HomeEmergency -- Using Audio to Find and Respond to Emergencies in the Home
James F. Mullen, Dhruva Kumar, Xuewei Qi +4
In the United States alone accidental home deaths exceed 128,000 per year. Our work aims to enable home robots who respond to emergency scenarios in the home, preventing injuries a…
ET-Former: Efficient Triplane Deformable Attention for 3D Semantic Scene Completion From Monocular Camera
Jing Liang, He Yin, Xuewei Qi +4
We introduce ET-Former, a novel end-to-end algorithm for semantic scene completion using a single monocular camera. Our approach generates a semantic occupancy map from single RGB…