works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.SD2026

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

Liting Gao, Yonggang Zhu, Yaru Chen +5

The paper introduces RFM-Editing 2, a text‑guided audio editing system that uses rectified flow matching and a two‑stage diffusion transformer with a coarse‑to‑fine attention schem…

cs.SD2026

Grammar-Guided Hierarchical Parsing for Long-form Audio Activity Recognition

Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1

Long-form audio exhibits an inherent hierarchy: fine-grained events form sub-activities, which in turn constitute higher-level activities. Prior work often models these levels sepa…

cs.SD2026

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

Yonggang Zhu, Liting Gao, Aidong Men +1

Contrastive Language-Audio Pretraining (CLAP) models are widely used for audio understanding and support modality-agnostic condition swapping in many zero-shot applications. Howeve…

cs.CV2026

TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation

Xinran Liu, Diptesh Kanojia, Wenwu Wang +1

Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making i…

cs.SD2026

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing

Liting Gao, Yi Yuan, Yaru Chen +5

Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the targ…

cs.CV2026

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips

Mengtian Li, Kunyan Dai, Yi Ding +4

Foley art plays a pivotal role in enhancing immersive auditory experiences in film, yet manual creation of spatio-temporally aligned audio remains labor-intensive. We propose Foley…