papers

Publications (13)

cs.CV2026

PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving

Pin Tang, Guoqing Wang, Xiangxuan Ren +4

Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization in complex autonomous drivi…

cs.CV2026

Grounding Everything in Tokens for Multimodal Large Language Models

Xiangxuan Ren, Zhongdao Wang, Liping Hou +3

Multimodal large language models (MLLMs) have made significant advancements in vision understanding and reasoning. However, the autoregressive Transformer architecture used by MLLM…

cs.CV2026

Learning Vision-Language-Action World Models for Autonomous Driving

Guoqing Wang, Pin Tang, Xiangxuan Ren +3

Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and control within a unified mult…

cs.CV2026

Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving

Guoqing Wang, Pin Tang, Xiangxuan Ren +2

Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent state-of-the-art frameworks achieve promising…

cs.CV2024

OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving

Guoqing Wang, Zhongdao Wang, Pin Tang +4

Existing solutions for 3D semantic occupancy prediction typically treat the task as a one-shot 3D voxel-wise segmentation perception problem. These discriminative methods focus on…

cs.CV2026

SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction

Pin Tang, Zhongdao Wang, Guoqing Wang +2

Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty space, while BEV and TPV proje…

cs.CV2024

SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction

Pin Tang, Zhongdao Wang, Guoqing Wang +4

Vision-based perception for autonomous driving requires an explicit modeling of a 3D space, where 2D latent representations are mapped and subsequent 3D operators are applied. Howe…

eess.IV2020

DSU-net: Dense SegU-net for automatic head-and-neck tumor segmentation in MR images

Pin Tang, Chen Zu, Mei Hong +7

Precise and accurate segmentation of the most common head-and-neck tumor, nasopharyngeal carcinoma (NPC), in MRI sheds light on treatment and regulatory decisions making. However,…

cs.CV2025

Enhancing Sampling Protocol for Point Cloud Classification Against Corruptions

Chongshou Li, Pin Tang, Xinke Li +2

Established sampling protocols for 3D point cloud learning, such as Farthest Point Sampling (FPS) and Fixed Sample Size (FSS), have long been relied upon. However, real-world data…

cs.CV2024

VEON: Vocabulary-Enhanced Occupancy Prediction

Jilai Zheng, Pin Tang, Zhongdao Wang +4

Perceiving the world as 3D occupancy supports embodied agents to avoid collision with any types of obstacle. While open-vocabulary image understanding has prospered recently, how t…

cs.CL2020

Recognizing Chinese Judicial Named Entity using BiLSTM-CRF

Pin Tang, Pinli Yang, Yuang Shi +3

Named entity recognition (NER) plays an essential role in natural language processing systems. Judicial NER is a fundamental component of judicial information retrieval, entity rel…

cs.CV2026

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation

Haoyu Zhao, Zihao Zhang, Jiaxi Gu +10

Camera-controllable video generation aims to synthesize videos with flexible and physically plausible camera movements. However, existing methods either provide imprecise camera co…

cs.CV2025

LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation

Xiangxuan Ren, Zhongdao Wang, Pin Tang +3

3D object detection is fundamental for safe and robust intelligent transportation systems. Current multi-modal 3D object detectors often rely on complex architectures and training…