5 papers · 1 filter
Egocentric Vision Language Planning
Zhirui Fang, Ming Yang, Weishuai Zeng +5
We explore leveraging large multi-modal models (LMMs) and text2image models to build a more general embodied agent. LMMs excel in planning long-horizon tasks over symbolic abstract…
Crowd-Sourced NeRF: Collecting Data from Production Vehicles for 3D Street View Reconstruction
Tong Qin, Changze Li, Haoyang Ye +4
Recently, Neural Radiance Fields (NeRF) achieved impressive results in novel view synthesis. Block-NeRF showed the capability of leveraging NeRF to build large city-scale models. F…
Unpaired Multi-view Clustering via Reliable View Guidance
Like Xin, Wanqi Yang, Lei Wang +1
This paper focuses on unpaired multi-view clustering (UMC), a challenging problem where paired observed samples are unavailable across multiple views. The goal is to perform effect…
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
Qingpei Guo, Furong Xu, Hanxiao Zhang +6
Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and…
SAM4UDASS: When SAM Meets Unsupervised Domain Adaptive Semantic Segmentation in Intelligent Vehicles
Weihao Yan, Yeqiang Qian, Xingyuan Chen +3
Semantic segmentation plays a critical role in enabling intelligent vehicles to comprehend their surrounding environments. However, deep learning-based methods usually perform poor…