9 papers
Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models
Yuting Huang, Leilei Ding, Zhipeng Tang +7
While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulat…
FARMER: Flow AutoRegressive Transformer over Pixels
Guangting Zheng, Qinyu Zhao, Tao Yang +6
Directly modeling the explicit likelihood of the raw data distribution is key topic in the machine learning area, which achieves the scaling successes in Large Language Models by a…
VLMPlanner: Integrating Visual Language Models with Motion Planning
Zhipeng Tang, Sha Zhang, Jiajun Deng +5
Integrating large language models (LLMs) into autonomous driving motion planning has recently emerged as a promising direction, offering enhanced interpretability, better controlla…
PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum
Shiqi Zhang, Sha Zhang, Jiajun Deng +3
Existing open-vocabulary 3D semantic segmentation methods typically supervise 3D segmentation models by merging text-aligned features (e.g., CLIP) extracted from multi-view images…
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
Guangting Zheng, Yehao Li, Yingwei Pan +4
Autoregressive models have emerged as a powerful generative paradigm for visual generation. The current de-facto standard of next token prediction commonly operates over a single-s…
SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
Yu Sheng, Jiajun Deng, Xinran Zhang +4
A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate se…