activity
20212024
most citedMM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

80 citations · 153 across the 16 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS20241 cited

Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Naoyuki Kanda, Xiaofei Wang, Sefik Emre Eskimez +12

Laughter is one of the most expressive and natural aspects of human speech, conveying emotions, social cues, and humor. However, most text-to-speech (TTS) systems lack the ability…

eess.AS20231 cited

Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition

Shaoshi Ling, Yuxuan Hu, Shuangbei Qian +5

Most end-to-end (E2E) speech recognition models are composed of encoder and decoder blocks that perform acoustic and language modeling functions. Pretrained large language models (…

eess.AS20231 cited

Adapting Multi-Lingual ASR Models for Handling Multiple Talkers

Chenda Li, Yao Qian, Zhuo Chen +5

State-of-the-art large-scale universal speech models (USMs) show a decent automatic speech recognition (ASR) performance across multiple domains and languages. However, it remains…

eess.AS2023

Code-Switching Text Generation and Injection in Mandarin-English ASR

Haibin Yu, Yuxuan Hu, Yao Qian +7

Code-switching speech refers to a means of expression by mixing two or more languages within a single utterance. Automatic Speech Recognition (ASR) with End-to-End (E2E) modeling f…

eess.AS2023

Target Sound Extraction with Variable Cross-modality Clues

Chenda Li, Yao Qian, Zhuo Chen +5

Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture o…