2 citations · 6 across the 17 of their papers we have counts for
7 papers · 1 filter
AdaTok: Self-Budgeting Image Tokenization with Quality-Preserving Dynamic Tokens
Xiaocheng Lu, Yuxi Chen, Jie Zhang +5
Image tokenizers, from 2D grids to recent 1D sequences, typically encode every image with the same fixed number of tokens. Yet visual complexity is highly heterogeneous, so a unifo…
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
Junzhe Chen, Siyuan Meng, Yuxi Chen +3
Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored.…
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
Jiawei Yang, Ziyu Chen, Yurong You +7
We present Flex, an efficient and effective scene encoder that addresses the computational bottleneck of processing high-volume multi-camera data in end-to-end autonomous driving.…
Latent Chain-of-Thought World Modeling for End-to-End Driving
Shuhan Tan, Kashyap Chitta, Yuxiao Chen +8
Recent Vision-Language-Action (VLA) models for autonomous driving explore inference-time reasoning as a way to improve driving performance and safety in challenging scenarios. Most…
DreamDrive: Generative 4D Scene Modeling from Street View Images
Jiageng Mao, Boyi Li, Boris Ivanovic +7
Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based…
STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes
Jiawei Yang, Jiahui Huang, Yuxiao Chen +10
We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often…