collaborators

5 papers

cs.CV2026

SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion

Xinyu Chen, Yuyi Qian, Jiang Lin +9

Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hin…

cs.MM2026

CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration

Tianyidan Xie, Zhentao Huang, Mingjie Wang +4

Automated movie creation requires coordinating multiple characters, modalities, and narrative elements across extended sequences -- a challenge that existing end-to-end approaches…

cs.CV2026

PhysLayer: Language-Guided Layered Animation with Depth-Aware Physics

Tianyidan Xie, Zhentao Huang, Mingjie Wang +4

Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior approaches have incorporated phys…

cs.CV2025

Region-to-Region: Enhancing Generative Image Harmonization with Adaptive Regional Injection

Zhiqiu Zhang, Dongqi Fan, Mingjie Wang +3

The goal of image harmonization is to adjust the foreground in a composite image to achieve visual consistency with the background. Recently, latent diffusion model (LDM) are appli…

cs.CV2025

GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning

Jun Wang, Hao Ruan, Liangjian Wen +2

Visual gaze estimation, with its wide-ranging application scenarios, has garnered increasing attention within the research community. Although existing approaches infer gaze solely…