3 citations · 4 across the 3 of their papers we have counts for
5 papers
OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models
Jinze Bai, Rui Men, Hao Yang +15
Generalist models, which are capable of performing diverse multi-modal tasks in a task-agnostic way within a single model, have been explored recently. Being, hopefully, an alterna…
MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition
Xiaohuan Zhou, Jiaming Wang, Zeyu Cui +4
In this paper, we propose a novel multi-modal multi-task encoder-decoder pre-training framework (MMSpeech) for Mandarin automatic speech recognition (ASR), which employs both unlab…
Contextual Expressive Text-to-Speech
Jianhong Tu, Zeyu Cui, Xiaohuan Zhou +4
The goal of expressive Text-to-speech (TTS) is to synthesize natural speech with desired content, prosody, emotion, or timbre, in high expressiveness. Most of previous studies atte…
Speech2Slot: An End-to-End Knowledge-based Slot Filling from Speech
Pengwei Wang, Xin Ye, Xiaohuan Zhou +2
In contrast to conventional pipeline Spoken Language Understanding (SLU) which consists of automatic speech recognition (ASR) and natural language understanding (NLU), end-to-end S…
xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems
Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang +3
Combinatorial features are essential for the success of many commercial models. Manually crafting these features usually comes with high cost due to the variety, volume and velocit…