4 papers
Point2Insert: Video Object Insertion via Sparse Point Guidance
Yu Zhou, Xiaoyan Yang, Bojia Zi +6
This paper introduces Point2Insert, a sparse-point-based framework for flexible and user-friendly object insertion in videos, motivated by the growing popularity of accurate, low-e…
TeleStyle: Content-Preserving Style Transfer in Images and Videos
Shiwen Zhang, Xiaoyan Yang, Bojia Zi +3
Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the i…
TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model
Yabo Chen, Yuanzhi Liang, Jiepeng Wang +24
World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent vi…
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
Chi Zhang, Jiepeng Wang, Youming Wang +5
We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to…