From the 1 of 7 linked papers with an AI index.
7 papers
ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models
Haojie Ren, Songrui Luo, Lingfeng Wang +6
The paper introduces ViCo3D, a framework that leverages vision foundation models to enrich LiDAR bird's-eye-view features for collaborative 3D object detection in V2X scenarios, ac…
GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation
Yuan Zhang, Shiqi Zhang, Yedong Shen +11
Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot emb…
CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks
Beibei Wang, Zijian Yu, Lu Zhang +8
Numerous roadside perception datasets have been introduced to propel advancements in autonomous driving and intelligent transportation systems research and development. However, it…
Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models
Yuting Huang, Leilei Ding, Zhipeng Tang +7
While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulat…
SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
Yu Sheng, Jiajun Deng, Xinran Zhang +4
A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate se…
\(X\)-evolve: Solution space evolution powered by large language models
Yi Zhai, Zhiqiang Wei, Ruohan Li +7
While combining large language models (LLMs) with evolutionary algorithms (EAs) shows promise for solving complex optimization problems, current approaches typically evolve individ…