activity
20232026
most citedMinMo: A Multimodal Large Language Model for Seamless Voice Interaction

3 citations · 4 across the 14 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2025

Online Audio-Visual Autoregressive Speaker Extraction

Zexu Pan, Wupeng Wang, Shengkui Zhao +4

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less ex…

eess.AS2025

Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

Zexu Pan, Shengkui Zhao, Tingting Wang +4

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. Ho…

eess.AS2025

UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook

Yidi Jiang, Qian Chen, Shengpeng Ji +6

The emergence of audio language models is empowered by neural audio codecs, which establish critical mappings between continuous waveforms and discrete tokens compatible with langu…

eess.AS2024

Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions

Kun Zhou, You Zhang, Dianwen Ng +3

Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of e…

eess.AS2024

Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis

Kun Zhou, Shengkui Zhao, Yukun Ma +7

Recent language model-based text-to-speech (TTS) frameworks demonstrate scalability and in-context learning capabilities. However, they suffer from robustness issues due to the acc…