34 citations · 34 across the 9 of their papers we have counts for
8 papers
Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
Tianyidan Xie, Shenyi Wang, Qiang Tang +7
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to…
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
Xinyu Chen, Yuyi Qian, Jiang Lin +9
Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hin…
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
Tianyidan Xie, Zhentao Huang, Mingjie Wang +4
Automated movie creation requires coordinating multiple characters, modalities, and narrative elements across extended sequences -- a challenge that existing end-to-end approaches…
PhysLayer: Language-Guided Layered Animation with Depth-Aware Physics
Tianyidan Xie, Zhentao Huang, Mingjie Wang +4
Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior approaches have incorporated phys…
Region-to-Region: Enhancing Generative Image Harmonization with Adaptive Regional Injection
Zhiqiu Zhang, Dongqi Fan, Mingjie Wang +3
The goal of image harmonization is to adjust the foreground in a composite image to achieve visual consistency with the background. Recently, latent diffusion model (LDM) are appli…
One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild
Dongqi Fan, Tao Chen, Mingjie Wang +5
Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they ofte…