collaborators

5 papers

cs.SD2025

PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching

Ali Vosoughi, Yongyi Zang, Qihui Yang +3

Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: t…

cs.LG2025

Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders

Nathan Paek, Yongyi Zang, Qihui Yang +1

While sparse autoencoders (SAEs) successfully extract interpretable features from language models, applying them to audio generation faces unique challenges: audio's dense nature r…

cs.SD2025

StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks

Jingyue Huang, Qihui Yang, Fei Yueh Chen +4

Existing pitch curve generators face two main challenges: they often neglect singer-specific expressiveness, reducing their ability to capture individual singing styles. And they a…

cs.SD2025

FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search

Qihui Yang, Randal Leistikow, Yongyi Zang

Virtual instrument generation requires maintaining consistent timbre across different pitches and velocities, a challenge that existing note-level models struggle to address. We pr…

cs.SD2025

WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling

Qihui Yang, Taylor Berg-Kirkpatrick, Julian McAuley +1

Despite rapid progress in end-to-end AI music generation, AI-driven modeling of professional Digital Signal Processing (DSP) workflows remains challenging. In particular, while the…