collaborators

5 papers

cs.CV2025

Object Isolated Attention for Consistent Story Visualization

Xiangyang Luo, Junhao Cheng, Yifan Xie +5

Open-ended story visualization is a challenging task that involves generating coherent image sequences from a given storyline. One of the main difficulties is maintaining character…

cs.SD2025

STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation

Tao Feng, Zhiyuan Zhao, Yifan Xie +4

We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that r…

cs.CV2025

UniSync: A Unified Framework for Audio-Visual Synchronization

Tao Feng, Yifan Xie, Xun Guan +4

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this cha…

cs.CV2025

CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis

Yifan Xie, Jingge Wang, Tao Feng +2

Colonoscopy is crucial for identifying adenomatous polyps and preventing colorectal cancer. However, developing robust models for polyp detection is challenging by the limited size…

cs.SD2024

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

Yifan Xie, Tao Feng, Xin Zhang +6

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attent…