collaborators

9 papers

cs.SD2026

DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion

Hanlin Zhang, Daxin Tan, Dehua Tao +5

Speech tokenizers are a key building block of fully discrete Speech LLMs. Existing tokenizers either prioritize semantic encoding, fuse semantic content with acoustic style insepar…

eess.AS2026

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

Hanlin Zhang, Daxin Tan, Dehua Tao +3

Instruction-guided speech editing requires a model to modify specified speech attributes while preserving unrelated characteristics. Despite rapid progress in Speech Large Language…

eess.AS2026

A Survey of Audio Reasoning in Multimodal Foundation Models

Zhihan Guo, Wenqian Cui, Guan-Ting Lin +8

Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from tho…

cs.CL2026

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

Wenqian Cui, Xiao-Hui Li, Daxin Tan +2

Speech large language models (SLMs) are typically built from text large language model (TLM) checkpoints, yet they still suffer from a substantial modality gap. Prior work has main…

eess.AS2026

Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models

Dehua Tao, Xuan Luo, Daxin Tan +5

While large-scale omni-models have demonstrated impressive capabilities across various modalities, their strong performance heavily relies on massive multimodal data and incurs sub…

cs.CL2026

PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs

Jing Xu, Jiaqi Wang, Daxin Tan +1

Although Large Language Models (LLMs) excel in many tasks, their application to Speech-to-Speech Translation (S2ST) is underexplored and hindered by data scarcity. To bridge this g…