4 papers
DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving
Hongbin Lin, Yiming Yang, Chaoda Zheng +7
In autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, tra…
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Zhihao Yuan, Shuyi Jiang, Chun-Mei Feng +4
Currently, utilizing large language models to understand the 3D world is becoming popular. Yet existing 3D-aware LLMs act as black boxes: they output bounding boxes or textual answ…
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
Hongbin Lin, Zilu Guo, Yifan Zhang +5
In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of…
PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
Zilu Guo, Hongbin Lin, Zhihao Yuan +6
3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and subopt…