3 papers
cs.CV2026
World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration
Ye Chen, Xuanhong Chen, Yupeng Zhu +23
The fundamental obstacle to industrial grade video generation is the lack of controllability: existing models treat video as a pixel distribution sampling problem, bypassing the ex…
cs.CV2026
ProxyImg: Towards Highly-Controllable Image Representation via Hierarchical Disentangled Proxy Embedding
Ye Chen, Yupeng Zhu, Xiongzhen Zhang +4
Prevailing image representation methods, including explicit representations such as raster images and Gaussian primitives, as well as implicit representations such as latent images…
cs.CV2026
Efficient Token Pruning for LLaDA-V
Zhewen Wan, Tianchen Song, Chen Lin +2
Diffusion-based large multimodal models, such as LLaDA-V, have demonstrated impressive capabilities in vision-language understanding and generation. However, their bidirectional at…