From the 1 of 26 linked papers with an AI index.
26 papers
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Hengyi Xie, Chenfei Yao, Xianjin Wu +7
TurboVLA is a vision-language-action model that directly maps visual observations and language instructions to robot actions, achieving real-time performance (32 Hz) on an RTX 4090…
Wat3R: Underwater 3D Geometry Learning without Annotations
Jiangwei Ren, Xingyu Jiang, Zijie Song +4
Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pion…
Towards Generalizable Robotic Manipulation in Dynamic Environments
Heng Fang, Shangru Li, Shuhan Wang +3
Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap primarily stems from a scarcity of d…
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
Xin Zhou, Dingkang Liang, Xiwu Chen +4
Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene gen…
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
Siyuan Liu, Chaoqun Zheng, Xin Zhou +3
Scene-level point cloud understanding remains challenging due to diverse geometries, imbalanced category distributions, and highly varied spatial layouts. Existing methods improve…
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
Yiran Guan, Sifan Tu, Dingkang Liang +6
Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni-modal large language models (OLLM) excel…