activity
20242026
collaborators

6 papers

cs.CV2026

SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding

Xingcheng Zhou, Mingyu Liu, Walter Zimmer +2

We present Scene-Graph Based Multi-Modal Traffic Agent (SGTA), a modular framework for traffic video understanding that combines structured scene graphs with multi-modal reasoning.…

cs.CV2026

Safety-Aligned 3D Object Detection: Single-Vehicle, Cooperative, and End-to-End Perspectives

Brian Hsuan-Cheng Liao, Chih-Hong Cheng, Hasan Esen +1

Perception plays a central role in connected and autonomous vehicles (CAVs), underpinning not only conventional modular driving stacks, but also cooperative perception systems and…

cs.CV2025

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

Xingcheng Zhou, Xuyuan Han, Feng Yang +3

We present OpenDriveVLA, a Vision Language Action model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially g…

cs.CV2025

Generative AI for Autonomous Driving: Frontiers and Opportunities

Yuping Wang, Shuo Xing, Cui Can +44

Generative Artificial Intelligence (GenAI) constitutes a transformative technological wave that reconfigures industries through its unparalleled capabilities for content creation,…

cs.CV2025

LiDAR-Guided Monocular 3D Object Detection for Long-Range Railway Monitoring

Raul David Dominguez Sanchez, Xavier Diaz Ortiz, Xingcheng Zhou +4

Railway systems, particularly in Germany, require high levels of automation to address legacy infrastructure challenges and increase train traffic safely. A key component of automa…

eess.IV2024

PointCompress3D: A Point Cloud Compression Framework for Roadside LiDARs in Intelligent Transportation Systems

Walter Zimmer, Ramandika Pranamulia, Xingcheng Zhou +2

In the context of Intelligent Transportation Systems (ITS), efficient data compression is crucial for managing large-scale point cloud data acquired by roadside LiDAR sensors. The…