most citedEnhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora

1 citations · 1 across the 11 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation

Daxin Tan, Dehua Tao, Chengxi Deng +2

Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models. Although this design enables s…

eess.AS2026

Speaker-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement

Hanlin Zhang, Daxin Tan, Dehua Tao +3

Semantic speech tokens should preserve linguistic content while suppressing speaker- and duration-dependent variation inherited from acoustic inputs. We propose Iterative Semantic…

eess.AS2026

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

Hanlin Zhang, Daxin Tan, Dehua Tao +3

Instruction-guided speech editing requires a model to modify specified speech attributes while preserving non-target characteristics. Despite rapid progress in Speech Large Languag…

eess.AS2026

A Survey of Audio Reasoning in Multimodal Foundation Models

Zhihan Guo, Wenqian Cui, Guan-Ting Lin +8

Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from tho…

eess.AS2026

Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models

Dehua Tao, Xuan Luo, Daxin Tan +5

While large-scale omni-models have demonstrated impressive capabilities across various modalities, their strong performance heavily relies on massive multimodal data and incurs sub…

eess.AS20241 cited

Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora

Jing Xu, Daxin Tan, Jiaqi Wang +1

While Large Language Models (LLMs) have shown potential in speech generation and recognition, their applications are mainly confined to monolingual scenarios, with limited explorat…