papers

Publications (23)

cs.CV2023

SRCN3D: Sparse R-CNN 3D for Compact Convolutional Multi-View 3D Object Detection and Tracking

Yining Shi, Jingyan Shen, Yifan Sun +5

Detection and tracking of moving objects is an essential component in environmental perception for autonomous driving. In the flourishing field of multi-view 3D camera-based detect…

cs.RO2026

4DLidarOpen: An Open 4D FMCW Lidar Dataset for Motion-Aware Autonomous Driving

Kane Qian, Xin Zhao, Yining Shi +7

We present 4DLidarOpen, a large-scale open multi-modal dataset for autonomous driving, centered on 4D frequency-modulated continuous-wave (FMCW) Lidar sensing. Unlike conventional…

cs.RO2025

FASIONAD++ : Integrating High-Level Instruction and Information Bottleneck in FAt-Slow fusION Systems for Enhanced Safety in Autonomous Driving with Adaptive Feedback

Kangan Qian, Ziang Luo, Sicong Jiang +16

Ensuring safe, comfortable, and efficient planning is crucial for autonomous driving systems. While end-to-end models trained on large datasets perform well in standard driving sce…

cs.CV2024

PanoSSC: Exploring Monocular Panoptic 3D Scene Reconstruction for Autonomous Driving

Yining Shi, Jiusi Li, Kun Jiang +4

Vision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomou…

cs.CV2025

CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates

Ankit Kumar Shaw, Kun Jiang, Tuopu Wen +5

The rapid growth of intelligent connected vehicles (ICVs) and integrated vehicle-road-cloud systems has increased the demand for accurate, real-time HD map updates. However, ensuri…

cs.CV2024

Grid-Centric Traffic Scenario Perception for Autonomous Driving: A Comprehensive Review

Yining Shi, Kun Jiang, Jiusi Li +5

Grid-centric perception is a crucial field for mobile robot perception and navigation. Nonetheless, grid-centric perception is less prevalent than object-centric perception as auto…

cs.CV2025

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

Yining Shi, Kun Jiang, Qiang Meng +6

World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspe…

cs.CV2024

StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential Equation

Yining Shi, Kun Jiang, Ke Wang +4

Predicting the future occupancy states of the surrounding environment is a vital task for autonomous driving. However, current best-performing single-modality methods or multi-moda…

cs.RO2022

Map Container: A Map-based Framework for Cooperative Perception

Kun Jiang, Yining Shi, Benny Wijaya +4

The idea of cooperative perception is to benefit from shared perception data between multiple vehicles and overcome the limitations of on-board sensors on single vehicle. However,…

cs.CV2025

DriveCamSim: Generalizable Camera Simulation via Explicit Camera Modeling for Autonomous Driving

Wenchao Sun, Xuewu Lin, Keyu Chen +4

Camera sensor simulation serves as a critical role for autonomous driving (AD), e.g. evaluating vision-based AD algorithms. While existing approaches have leveraged generative mode…

cs.LG2021

HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework

Xupeng Miao, Hailin Zhang, Yining Shi +4

Embedding models have been an effective learning paradigm for high-dimensional data. However, one open issue of embedding models is that their representations (latent factors) ofte…

cs.CV2025

POD: Predictive Object Detection with Single-Frame FMCW LiDAR Point Cloud

Yining Shi, Kun Jiang, Xin Zhao +5

LiDAR-based 3D object detection is a fundamental task in the field of autonomous driving. This paper explores the unique advantage of Frequency Modulated Continuous Wave (FMCW) LiD…

cs.CV2025

How Cars Move: Analyzing Driving Dynamics for Safer Urban Traffic

Kangan Qian, Jinyu Miao, Xinyu Jiao +6

Understanding the spatial dynamics of cars within urban systems is essential for optimizing infrastructure management and resource allocation. Recent empirical approaches for analy…

cs.CV2026

SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving

Wenchao Sun, Xuewu Lin, Keyu Chen +4

End-to-end multi-modal planning has been widely adopted to model the uncertainty of driving behavior, typically by scoring candidate trajectories and selecting the optimal one. Exi…

cs.RO2025

AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving

Kangan Qian, Sicong Jiang, Yang Zhong +18

Vision-Language Models (VLMs) show promise for autonomous driving, yet their struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate…

cs.LG2024

From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers

Shaoxiong Duan, Yining Shi, Wei Xu

In this paper, we investigate the inherent capabilities of transformer models in learning arithmetic algorithms, such as addition and parity. Through experiments and attention anal…

cs.CV2022

Bridging the View Disparity Between Radar and Camera Features for Multi-modal Fusion 3D Object Detection

Taohua Zhou, Yining Shi, Junjie Chen +3

Environmental perception with the multi-modal fusion of radar and camera is crucial in autonomous driving to increase accuracy, completeness, and robustness. This paper focuses on…

cs.RO2024

FASIONAD : FAst and Slow FusION Thinking Systems for Human-Like Autonomous Driving with Adaptive Feedback

Kangan Qian, Zhikun Ma, Yangfan He +13

Ensuring safe, comfortable, and efficient navigation is a critical goal for autonomous driving systems. While end-to-end models trained on large-scale datasets excel in common driv…

cs.CV2025

LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction

Kangan Qian, Jinyu Miao, Ziang Luo +7

Accurate and reliable spatial and motion information plays a pivotal role in autonomous driving systems. However, object-level perception models struggle with handling open scenari…

cs.LG2025

TileLang: A Composable Tiled Programming Model for AI Systems

Lei Wang, Yu Cheng, Yining Shi +8

Modern AI workloads rely heavily on optimized computing kernels for both training and inference. These AI kernels follow well-defined data-flow patterns, such as moving tiles betwe…

cs.CV2025

EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Driving

Yining Shi, Kun Jiang, Jinyu Miao +8

3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computatio…

cs.CV2024

SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation

Wenchao Sun, Xuewu Lin, Yining Shi +3

The well-established modular autonomous driving system is decoupled into different standalone tasks, e.g. perception, prediction and planning, suffering from information loss and e…

cs.SD2024

Soundify: Matching Sound Effects to Video

David Chuan-En Lin, Anastasis Germanidis, Cristóbal Valenzuela +2

In the art of video editing, sound helps add character to an object and immerse the viewer within a space. Through formative interviews with professional editors (N=10), we found t…