activity
20242026
collaborators

15 papers

cs.RO2026

Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future

Tianshuai Hu, Xiaolu Liu, Song Wang +17

Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-ta…

cs.CV2025

Learning to Remove Lens Flare in Event Camera

Haiqian Han, Lingdong Kong, Jianing Li +7

Event cameras have the potential to revolutionize vision systems with their high temporal resolution and dynamic range, yet they remain susceptible to lens flare, a fundamental opt…

cs.CV2025

3EED: Ground Everything Everywhere in 3D

Rong Li, Yuhao Dong, Tianshuai Hu +7

Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, si…

cs.CV2025

Learning to Generate 4D LiDAR Sequences

Ao Liang, Youquan Liu, Yu Yang +5

While generative world models have advanced video and occupancy-based data synthesis, LiDAR generation remains underexplored despite its importance for accurate 3D perception. Exte…

cs.CV2025

Visual Grounding from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang +6

Event cameras capture changes in brightness with microsecond precision and remain reliable under motion blur and challenging illumination, offering clear advantages for modeling hi…

cs.CV2025

La La LiDAR: Large-Scale Layout Generation from LiDAR Data

Youquan Liu, Lingdong Kong, Weidong Yang +5

Controllable generation of realistic LiDAR scenes is crucial for applications such as autonomous driving and robotics. While recent diffusion-based models achieve high-fidelity LiD…