activity
20242026
collaborators

7 papers

cs.CV2026

In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion

Lingxiao Yang, Liu Liu, Moran Li +4

Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean…

cs.CV2026

SteerVTE: Seamless Video Text Editing with Style and Glyph Control

Kai Zeng, Moran Li, Zhengwei Wang +6

Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain,…

cs.CV2026

TextSculptor: Training and Benchmarking Scene Text Editing

Yiheng Lin, Siyu Jiao, Xiaohan Lan +12

Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editi…

cs.CV2025

Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses

Yuji Wang, Moran Li, Xiaobin Hu +7

Current video generation models struggle with identity preservation under large face poses, primarily facing two challenges: the difficulty in exploring an effective mechanism to i…

cs.CV2025

Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations

Yuji Wang, Moran Li, Xiaobin Hu +7

Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. Howe…

cs.CV2025

StrandDesigner: Towards Practical Strand Generation with Sketch Guidance

Na Zhang, Moran Li, Chengming Xu +6

Realistic hair strand generation is crucial for applications like computer graphics and virtual reality. While diffusion models can generate hairstyles from text or images, these i…