28 papers
Unified Video Dense Prediction from Disjoint Data
Yihong Sun, Seoung Wug Oh, Jiahui Huang +2
Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, doma…
MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching
Jiahui Huang, Yasi Zhang, Tianyu Chen +6
Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality…
Déjà View: Looping Transformers for Multi-View 3D Reconstruction
Alessandro Burzio, Tobias Fischer, Sven Elflein +9
Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emergi…
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
Jinghui Lu, Jiayi Guan, Zhijian Huang +47
Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes a latency cost that is…
Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation
Tianshi Cao, Jiawei Ren, Yuxuan Zhang +12
Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before real-world deployment. Neural s…
TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens
Jiawei Ren, Michal Jan Tyszkiewicz, Jiahui Huang +1
In this work, we revisit several key design choices of modern Transformer-based approaches for feed-forward 3D Gaussian Splatting (3DGS) prediction. We argue that the common practi…