activity
20172026
most citedBIT: Biologically Inspired Tracker

47 citations · 67 across the 30 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2026

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

Jialong Mai, Xiaofen Xing, Xiangmin Xu

Fine-grained local timing control is still absent from modern text-to-speech systems: existing approaches typically provide only utterance-level duration or global speaking-rate co…

cs.SD2025

BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis

Jingyuan Xing, Mingru Yang, Zhipeng Li +2

Autoregressive (AR) frameworks have recently achieved remarkable progress in zero-shot text-to-speech (TTS) by leveraging discrete speech tokens and large language model techniques…

cs.SD2025

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

Jialong Mai, Jinxin Ji, Xiaofen Xing +4

Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as…

cs.SD2025

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models

Yuanbo Fang, Haoze Sun, Jun Liu +5

End-to-end speech large language models ((LLMs)) extend the capabilities of text-based models to directly process and generate audio tokens. However, this often leads to a decline…

cs.SD2023

Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition

Weidong Chen, Xiaofen Xing, Peihao Chen +1

This paper presents a paradigm that adapts general large-scale pretrained models (PTMs) to speech emotion recognition task. Although PTMs shed new light on artificial general intel…

cs.SD20231 cited

DWFormer: Dynamic Window transFormer for Speech Emotion Recognition

Shuaiqi Chen, Xiaofen Xing, Weibin Zhang +2

Speech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreov…