8 papers
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement
Jingbin Hu, Haoyu Zhang, Dake Guo +10
Large Language Models (LLMs) have advanced audio generation through discrete representation learning. However, most existing neural codecs focus on speech and emphasize reconstruct…
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
Wenjie Tian, Mingchen Shao, Bingshen Mu +8
Audio-visual speech recognition (AVSR) is an extension of ASR that incorporates visual signals. Current AVSR approaches primarily focus on lip motion, largely overlooking rich cont…
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
Wenjie Tian, Zhixian Zhao, Jingbin Hu +4
The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However,…
VoiceSculptor: Your Voice, Designed By You
Jingbin Hu, Huakang Chen, Linhan Ma +19
Despite rapid progress in text-to-speech (TTS), open-source systems still lack truly instruction-following, fine-grained control over core speech attributes (e.g., pitch, speaking…
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
Chengyou Wang, Mingchen Shao, Jingbin Hu +11
Speech processing for low-resource dialects remains a fundamental challenge in developing inclusive and robust speech technologies. Despite its linguistic significance and large sp…
HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
Ziyu Zhang, Hanzhao Li, Jingbin Hu +2
Controllable speech synthesis refers to the precise control of speaking style by manipulating specific prosodic and paralinguistic attributes, such as gender, volume, speech rate,…