collaborators

5 papers

cs.CV2026

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

Jialun Liu, Tian Li, Xiao Cao +20

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and r…

cs.CV2026

Point2Insert: Video Object Insertion via Sparse Point Guidance

Yu Zhou, Xiaoyan Yang, Bojia Zi +6

This paper introduces Point2Insert, a sparse-point-based framework for flexible and user-friendly object insertion in videos, motivated by the growing popularity of accurate, low-e…

cs.CV2026

TeleStyle: Content-Preserving Style Transfer in Images and Videos

Shiwen Zhang, Xiaoyan Yang, Bojia Zi +3

Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the i…

cs.CV2025

TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model

Yabo Chen, Yuanzhi Liang, Jiepeng Wang +24

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent vi…

cs.CV2025

UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation

Chi Zhang, Jiepeng Wang, Youming Wang +5

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to…