activity
20242026
collaborators

10 papers

cs.LG2026

Ratio-Variance Regularized Policy Optimization

Yu Luo, Shuo Han, Yihan Hu +5

Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscriminately truncating high-return…

eess.AS2026

StepAudio 2.5 Technical Report

Bin Lin, Bo Zhao, Boyong Wu +98

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…

cs.CV2025

Audio Frequency-Time Dual Domain Evaluation on Depression Diagnosis

Yu Luo, Nan Huang, Sophie Yu +5

Depression, as a typical mental disorder, has become a prevalent issue significantly impacting public health. However, the prevention and treatment of depression still face multipl…

cs.CL2025

Step-Audio 2 Technical Report

Boyong Wu, Chao Yan, Chen Hu +106

This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent…

cs.LG2025

Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding

StepFun, :, Bin Wang +195

Large language models (LLMs) face low hardware efficiency during decoding, especially for long-context reasoning tasks. This paper introduces Step-3, a 321B-parameter VLM with hard…

eess.AS2025

ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality

Yu-Xiang Luo, Yi-Cheng Lin, Ming-To Chuang +9

Despite extensive research on toxic speech detection in text, a critical gap remains in handling spoken Mandarin audio. The lack of annotated datasets that capture the unique proso…