activity
20232025
most citedA Quantitative Approach to Understand Self-Supervised Models as Cross-lingual Feature Extractors

1 citations · 2 across the 6 of their papers we have counts for

collaborators

10 papers

cs.AI2025

Step-Audio-R1 Technical Report

Fei Tian, Xiangyu Tony Zhang, Yuxin Zhang +14

Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon…

cs.CL2025

Step-Audio-EditX Technical Report

Chao Yan, Boyong Wu, Peng Yang +12

We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguisti…

eess.AS2025

MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models

Yayue Deng, Guoqiang Hu, Haiyang Sun +6

Spoken Dialogue Models (SDMs) have advanced rapidly, yet their ability to sustain genuinely interactive multi-turn conversations remains underexplored, as most benchmarks focus on…

cs.SD2025

Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech

Xinlei Niu, Jianbo Ma, Dylan Harper-Harris +3

The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on…

cs.CL20251 cited

Step-Audio 2 Technical Report

Boyong Wu, Chao Yan, Chen Hu +106

This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent…

eess.AS2025

Long-Context Modeling Networks for Monaural Speech Enhancement: A Comparative Study

Qiquan Zhang, Moran Chen, Zeyang Song +3

Advanced long-context modeling backbone networks, such as Transformer, Conformer, and Mamba, have demonstrated state-of-the-art performance in speech enhancement. However, a system…