collaborators

9 papers

cs.SD2026

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

Rongshen He, Xinyu Liang, Dekun Chen +3

Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentenc…

eess.AS2026

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

Haolong Zheng, Yuanzhuo Hu, Xinyu Liang +7

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and l…

cs.SD2026

FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates

Jiaqi Li, Chaoren Wang, Xiaohai Tian +9

Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying informati…

eess.AS2026

DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning

Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3

We introduce DNSMOS-C, a compact end-to-end speech quality assessment model that extends the DNSMOS Pro framework by integrating a MOS-guided triplet-based contrastive loss. Applie…

eess.AS2026

SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment

Fengyuan Cao, Xinyu Liang, Fredrik Cumlin +4

Designing a speech quality assessment (SQA) system for estimating mean-opinion-score (MOS) of multi-rate speech with varying sampling frequency (16-48 kHz) is a challenging task. T…

eess.AS2025

Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech

Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3

Self-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each c…