collaborators

8 papers

cs.CV2025

Fine-Grained Instruction-Guided Graph Reasoning for Vision-and-Language Navigation

Yaohua Liu, Xinyuan Song, Yunfu Deng +3

Vision-and-Language Navigation (VLN) requires an embodied agent to traverse complex environments by following natural language instructions, demanding accurate alignment between vi…

cs.CV2025

Universal Visuo-Tactile Video Understanding for Embodied Interaction

Yifan Xie, Mingyang Li, Shoujie Li +5

Tactile perception is essential for embodied agents to understand physical attributes of objects that cannot be determined through visual inspection alone. While existing approache…

cs.CV2025

Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning

Yifan Xie, Fei Ma, Yi Bin +2

Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the signific…

cs.CV2025

MuseFace: Text-driven Face Editing via Diffusion-based Mask Generation Approach

Xin Zhang, Siting Huang, Xiangyang Luo +5

Face editing modifies the appearance of face, which plays a key role in customization and enhancement of personal images. Although much work have achieved remarkable success in tex…

cs.CV2025

Object Isolated Attention for Consistent Story Visualization

Xiangyang Luo, Junhao Cheng, Yifan Xie +5

Open-ended story visualization is a challenging task that involves generating coherent image sequences from a given storyline. One of the main difficulties is maintaining character…

cs.SD2025

STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation

Tao Feng, Zhiyuan Zhao, Yifan Xie +4

We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that r…