collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20253 cited

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

Wei Chow, Jiageng Mao, Boyi Li +3

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. Whi…

cs.CV2025

DreamDrive: Generative 4D Scene Modeling from Street View Images

Jiageng Mao, Boyi Li, Boris Ivanovic +7

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based…

cs.CV20241 cited

STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes

Jiawei Yang, Jiahui Huang, Yuxiao Chen +10

We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often…

cs.CV2024

Promptable Closed-loop Traffic Simulation

Shuhan Tan, Boris Ivanovic, Yuxiao Chen +5

Simulation stands as a cornerstone for safe and efficient autonomous driving development. At its core a simulation system ought to produce realistic, reactive, and controllable tra…

cs.CV2024

Language-Image Models with 3D Understanding

Jang Hyun Cho, Boris Ivanovic, Yulong Cao +8

Multi-modal large language models (MLLMs) have shown incredible capabilities in a variety of 2D vision and language tasks. We extend MLLMs' perceptual capabilities to ground and re…