activity
20242026
collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen +27

MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped trans…

cs.SD2026

MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions

Kexin Huang, Liwei Fan, Botian Jiang +11

Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, person…

cs.SD2026

MOSS-TTSD: Text to Spoken Dialogue Generation

Yuqian Zhang, Donghua Yu, Zhengyuan Lin +15

Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance t…

cs.SD2025

HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection

Qing Wang, Ya Jiang, Hang Chen +3

This work presents HDA-SELD, a unified framework that combines hierarchical cross-modal distillation (HCMD) and multi-level data augmentation to address low-resource audio-visual (…

cs.SD2025

An Experimental Study on Joint Modeling for Sound Event Localization and Detection with Source Distance Estimation

Yuxuan Dong, Qing Wang, Hengyi Hong +2

In traditional sound event localization and detection (SELD) tasks, the focus is typically on sound event detection (SED) and direction-of-arrival (DOA) estimation, but they fall s…