collaborators

6 papers

cs.SD2025

Efficient Vocal Source Separation Through Windowed Sink Attention

Christodoulos Benetatos, Yongyi Zang, Randal Leistikow

State-of-the-art vocal separation models like Mel-Band-Roformer rely on full temporal self-attention mechanisms, where each temporal frame interacts with every other frames. This i…

cs.SD2025

PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching

Ali Vosoughi, Yongyi Zang, Qihui Yang +3

Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: t…

cs.LG2025

Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders

Nathan Paek, Yongyi Zang, Qihui Yang +1

While sparse autoencoders (SAEs) successfully extract interpretable features from language models, applying them to audio generation faces unique challenges: audio's dense nature r…

cs.SD2025

StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks

Jingyue Huang, Qihui Yang, Fei Yueh Chen +4

Existing pitch curve generators face two main challenges: they often neglect singer-specific expressiveness, reducing their ability to capture individual singing styles. And they a…

cs.SD2025

FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search

Qihui Yang, Randal Leistikow, Yongyi Zang

Virtual instrument generation requires maintaining consistent timbre across different pitches and velocities, a challenge that existing note-level models struggle to address. We pr…

cs.SD2025

Smule Renaissance Small: Efficient General-Purpose Vocal Restoration

Yongyi Zang, Chris Manchester, David Young +9

Vocal recordings on consumer devices commonly suffer from multiple concurrent degradations: noise, reverberation, band-limiting, and clipping. We present Smule Renaissance Small (S…