1 citations · 1 across the 2 of their papers we have counts for
14 papers
Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation
Yueting Zhu, Yuehao Song, Kaicheng Zhang +5
Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous video stream. However, streaming…
ACEsplat: Accelerated 3D Gaussian Scene Regression via RGB and Poses Only
Mingkai Liu, Haohua Que, Dikai Fan +7
Per-scene 3D Gaussian Splatting (3DGS) enables high-fidelity rendering, but practical robotic and AR scene capture pipelines often depend on external geometric initialization (e.g.…
SenseExpo: Spatial Exploration and Navigation via Scene Estimation from Expeditious Predictive Operators
Haojia Gao, Haohua Que, Mingkai Liu +9
We present \textbf{SenseExpo}, a lightweight single-robot exploration framework that integrates a compact map prediction network into a frontier-based strategy. SenseExpo addresses…
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
Bo Jiang, Shaoyu Chen, Hao Gao +4
Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of planning make it challenging. Existin…
RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework
Hao Gao, Shaoyu Chen, Yifan Zhu +4
High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interactions. Although diffusion-ba…
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
Hongyuan Tao, Bencheng Liao, Shaoyu Chen +4
Vision-Language Models (VLMs) are increasingly tasked with ultra-long multimodal understanding. While linear architectures offer constant computation and memory footprints, they of…