activity
20242026
collaborators

6 papers

cs.GR2026

MoSAT: Human Motion Generation from Spatial Audio and Textual Description

Shuyang Xu, Zhiyang Dou, Yiduo Hao +8

Human motion is shaped by both external acoustic events and behavioral intent: spatial audio conveys environmental cues that elicit or guide a response, while text specifies the de…

cs.CV2026

Next-Scale Autoregressive Models for Text-to-Motion Generation

Zhiwei Zheng, Shibo Jin, Lingjie Liu +1

Autoregressive (AR) models offer stable and efficient training, but standard next-token prediction is not well aligned with the temporal structure required for text-conditioned mot…

cs.SD2025

Building Audio-Visual Digital Twins with Smartphones

Zitong Lan, Yiwei Tang, Yuhan Wang +3

Digital twins today are almost entirely visual, overlooking acoustics-a core component of spatial realism and interaction. We introduce AV-Twin, the first practical system that con…

cs.SD2025

Resounding Acoustic Fields with Reciprocity

Zitong Lan, Yiduo Hao, Mingmin Zhao

Achieving immersive auditory experiences in virtual environments requires flexible sound modeling that supports dynamic source positions. In this paper, we introduce a task called…

cs.SD2025

Guiding Audio Editing with Audio Language Model

Zitong Lan, Yiduo Hao, Mingmin Zhao

Audio editing plays a central role in VR/AR immersion, virtual conferencing, sound design, and other interactive media. However, recent generative audio editing models depend on te…

cs.SD2024

Acoustic Volume Rendering for Neural Impulse Response Fields

Zitong Lan, Chenhao Zheng, Zhiwei Zheng +1

Realistic audio synthesis that captures accurate acoustic phenomena is essential for creating immersive experiences in virtual and augmented reality. Synthesizing the sound receive…