8 papers
PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
Yining Pan, Shijie Li, Yuchen Wu +2
This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve generalization under domain shifts…
Perception-Aware Multimodal Spatial Reasoning from Monocular Images
Yanchun Cheng, Rundong Wang, Xulei Yang +4
Spatial reasoning from monocular images is essential for autonomous driving, yet current Vision-Language Models (VLMs) still struggle with fine-grained geometric perception, partic…
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
Yongyi Su, Haojie Zhang, Shijie Li +11
Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…
DiffPCN: Latent Diffusion Model Based on Multi-view Depth Images for Point Cloud Completion
Zijun Li, Hongyu Yan, Shijie Li +4
Latent diffusion models (LDMs) have demonstrated remarkable generative capabilities across various low-level vision tasks. However, their potential for point cloud completion remai…
Zero-Shot 3D Visual Grounding from Vision-Language Models
Rong Li, Shijie Li, Lingdong Kong +2
3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as augmented reality and robotics.…
Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion
Yifan Liu, Xun Xu, Shijie Li +2
Multi-camera systems provide richer contextual information for industrial anomaly detection. However, traditional methods process each view independently, disregarding the compleme…