10 papers
VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation
Jiarui Hai, Karan Thakkar, Ke Chen +5
Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions. However, exi…
SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data
Bruno Santos Teixeira
Natural-language interfaces to enterprise data must translate underspecified requests into governed, executable behavior while controlling invalid queries, policy failures, cost, a…
DECAF: Dynamic Envelope Context-Aware Fusion for Speech-Envelope Reconstruction from EEG
Karan Thakkar, Mounya Elhilali
Reconstructing the speech audio envelope from scalp neural recordings (EEG) is a central task for decoding a listener's attentional focus in applications like neuro-steered hearing…
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
Xilin Jiang, Qiaolin Wang, Junkai Wu +30
Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand suc…
FlexSED: Towards Open-Vocabulary Sound Event Detection
Jiarui Hai, Helin Wang, Weizhe Guo +1
Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fund…
SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering
Jiarui Hai, Mounya Elhilali
Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up…