11 papers
MoWorld: A Flash World Model
Team Moxin, Deyi Ji, Tianrun Chen +37
The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…
Mean Flow Distillation: Robust and Stable Distillation for Flow Matching Models
An Zhao, Shengyuan Zhang, Zhongjian Sun +5
Flow Matching models have demonstrated strong performance across a wide range of generative tasks. However, their reliance on ODE-based iterative sampling incurs substantial comput…
4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
Ying Zang, Xuanyi Liu, Yidong Han +9
Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance s…
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
Xuanyi Liu, Chunan Yu, Deyi Ji +6
Reconstructing dense 3D geometry from continuous video streams requires stable inference under a constant memory budget. Existing frameworks primarily rely on a ``pure evict…
Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors
Ying Zang, Yidong Han, Chaotao Ding +8
Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences whe…
Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes
Chenye Meng, Zejian Li, Zhongni Liu +9
Post-training alignment of diffusion models relies on simplified signals, such as scalar rewards or binary preferences. This limits alignment with complex human expertise, which is…