collaborators

12 papers

cs.CV2026

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion

Bowen Cui, Weijie Wang, Zeyu Zhang +7

While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains challenging. Existing text-to-3D methods either decode discrete sh…

cs.MM2026

SPEED: One-Step Pixel Diffusion for High-quality Video Frame Interpolation

Zihao Zhang, Haoyu Zhao, Siqian Yang +3

Despite the success of diffusion models in Video Frame Interpolation (VFI), existing methods still suffer from two critical limitations. First, latent diffusion inevitably loses fi…

cs.CV2026

Latent Spatial Memory for Video World Models

Weijie Wang, Haoyu Zhao, Yifan Yang +7

Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computat…

cs.CV2026

CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping

Haoyu Zhao, Jiaxi Gu, Haoran Chen +11

Precise camera pose control is critical for video diffusion, yet maintaining geometric consistency remains a challenge. Existing methods that directly inject numerical camera param…

cs.RO2026

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

Zhiyuan Feng, Qixiu Li, Huizhi Liang +12

Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large…

cs.CV2026

DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation

Haoyu Zhao, Yuang Zhang, Junqi Cheng +5

Recent video generative models have demonstrated impressive visual fidelity, yet they often struggle with semantic, geometric, and identity consistency. In this paper, we propose a…