5 papers
TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment
Yu Meng, Xiangyang Luo, Letian Li +5
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generate…
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
Xiangyang Luo, Xiaozhe Xin, Tao Feng +3
Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite…
Human Motion Video Generation: A Survey
Haiwei Xue, Xiangyang Luo, Zhanghao Hu +12
Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars…
MuseFace: Text-driven Face Editing via Diffusion-based Mask Generation Approach
Xin Zhang, Siting Huang, Xiangyang Luo +5
Face editing modifies the appearance of face, which plays a key role in customization and enhancement of personal images. Although much work have achieved remarkable success in tex…
STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
Tao Feng, Zhiyuan Zhao, Yifan Xie +4
We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that r…