12 papers
MiniWorld: Democratizing the Training of Video World Models from Scratch
Yian Zhao, Ruochong Zheng, Hongcan Guo +3
Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions…
Continuous Latent Diffusion Language Model
Hongcan Guo, Qinyu Zhao, Yian Zhao +8
Large language models have achieved remarkable success under the autoregressive paradigm, yet high-quality text generation need not be tied to a fixed left-to-right order. Existing…
Video Generation with Predictive Latents
Yian Zhao, Feng Wang, Qiushan Guo +4
Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, improving training efficiency an…
Seedance 2.0: Advancing Video Generation for World Complexity
Team Seedance, De Chen, Liyang Chen +168
Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro…
SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection
Yifan Wang, Yian Zhao, Fanqi Pu +4
Existing monocular 3D detectors typically tame the pronounced nonlinear regression of 3D bounding box through decoupled prediction paradigm, which employs multiple branches to esti…
Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian Splatting
Yian Zhao, Rushi Ye, Ruochong Zheng +6
3D style transfer refers to the artistic stylization of 3D assets based on reference style images. Recently, 3DGS-based stylization methods have drawn considerable attention, prima…