From the 1 of 8 linked papers with an AI index.
8 papers
SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision
Hao Zhang, Yiwen Zhao, Yixuan Zhang +2
We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, a…
Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis
Mingyue Huo, Yuheng Zhang, Hao Zhang
The paper introduces an inference‑time polar projection method to diagnose how components of STFT‑domain speech enhancement affect automatic speech recognition, revealing that magn…
TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion
Zihan Zhang, Shreeram Suresh Chandra, Zongyang Du +5
Traditional emotional voice conversion (EVC) conditions generation on explicit target emotions like labels or references, defining the target affective state but omitting the direc…
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo +3
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, there is no consensus on whether audio-langu…
WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation
Mingda Lin, Lei Ding, Xinyue Zhou +6
While pre-trained models excel in specialized tasks, learning universal representations across diverse acoustic domains remains challenging. To address this, we propose WQ-Fusion,…
Covo-Audio Technical Report
Wenfu Wang, Chenxing Li, Liqiang Zhang +23
In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture…