activity
20242026
collaborators

7 papers

cs.CV2026

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

Bin Hu, Yanwen Ma, Jiehui Huang +14

Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering traje…

cs.AI2026

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond

Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin +47

As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck. Agents that ma…

cs.CV2026

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

Xu He, Haoxian Zhang, Hejia Chen +7

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing…

cs.CV2025

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation

Wenhui Song, Hanhui Li, Jiehui Huang +5

In this paper, we present LaVieID, a novel \underline{l}ocal \underline{a}utoregressive \underline{vi}d\underline{e}o diffusion framework designed to tackle the challenging \underl…

cs.CV2025

BridgeIV: Bridging Customized Image and Video Generation through Test-Time Autoregressive Identity Propagation

Panwen Hu, Jiehui Huang, Qiang Sun +1

Both zero-shot and tuning-based customized text-to-image (CT2I) generation have made significant progress for storytelling content creation. In contrast, research on customized tex…

cs.CV2025

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

Shiyue Zhang, Zheng Chong, Xi Lu +6

Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a prom…