5 papers
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
Jialun Liu, Tian Li, Xiao Cao +20
Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and r…
Point2Insert: Video Object Insertion via Sparse Point Guidance
Yu Zhou, Xiaoyan Yang, Bojia Zi +6
This paper introduces Point2Insert, a sparse-point-based framework for flexible and user-friendly object insertion in videos, motivated by the growing popularity of accurate, low-e…
TeleStyle: Content-Preserving Style Transfer in Images and Videos
Shiwen Zhang, Xiaoyan Yang, Bojia Zi +3
Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the i…
TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model
Yabo Chen, Yuanzhi Liang, Jiepeng Wang +24
World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent vi…
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
Chi Zhang, Jiepeng Wang, Youming Wang +5
We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to…