Publications (13)
PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving
Pin Tang, Guoqing Wang, Xiangxuan Ren +4
Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization in complex autonomous drivi…
Grounding Everything in Tokens for Multimodal Large Language Models
Xiangxuan Ren, Zhongdao Wang, Liping Hou +3
Multimodal large language models (MLLMs) have made significant advancements in vision understanding and reasoning. However, the autoregressive Transformer architecture used by MLLM…
Learning Vision-Language-Action World Models for Autonomous Driving
Guoqing Wang, Pin Tang, Xiangxuan Ren +3
Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and control within a unified mult…
Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving
Guoqing Wang, Pin Tang, Xiangxuan Ren +2
Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent state-of-the-art frameworks achieve promising…
OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving
Guoqing Wang, Zhongdao Wang, Pin Tang +4
Existing solutions for 3D semantic occupancy prediction typically treat the task as a one-shot 3D voxel-wise segmentation perception problem. These discriminative methods focus on…
SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction
Pin Tang, Zhongdao Wang, Guoqing Wang +2
Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty space, while BEV and TPV proje…
SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction
Pin Tang, Zhongdao Wang, Guoqing Wang +4
Vision-based perception for autonomous driving requires an explicit modeling of a 3D space, where 2D latent representations are mapped and subsequent 3D operators are applied. Howe…
DSU-net: Dense SegU-net for automatic head-and-neck tumor segmentation in MR images
Pin Tang, Chen Zu, Mei Hong +7
Precise and accurate segmentation of the most common head-and-neck tumor, nasopharyngeal carcinoma (NPC), in MRI sheds light on treatment and regulatory decisions making. However,…
Enhancing Sampling Protocol for Point Cloud Classification Against Corruptions
Chongshou Li, Pin Tang, Xinke Li +2
Established sampling protocols for 3D point cloud learning, such as Farthest Point Sampling (FPS) and Fixed Sample Size (FSS), have long been relied upon. However, real-world data…
VEON: Vocabulary-Enhanced Occupancy Prediction
Jilai Zheng, Pin Tang, Zhongdao Wang +4
Perceiving the world as 3D occupancy supports embodied agents to avoid collision with any types of obstacle. While open-vocabulary image understanding has prospered recently, how t…
Recognizing Chinese Judicial Named Entity using BiLSTM-CRF
Pin Tang, Pinli Yang, Yuang Shi +3
Named entity recognition (NER) plays an essential role in natural language processing systems. Judicial NER is a fundamental component of judicial information retrieval, entity rel…
CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
Haoyu Zhao, Zihao Zhang, Jiaxi Gu +10
Camera-controllable video generation aims to synthesize videos with flexible and physically plausible camera movements. However, existing methods either provide imprecise camera co…
LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation
Xiangxuan Ren, Zhongdao Wang, Pin Tang +3
3D object detection is fundamental for safe and robust intelligent transportation systems. Current multi-modal 3D object detectors often rely on complex architectures and training…