479 citations · 792 across the 34 of their papers we have counts for
47 papers
OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models
Jinze Bai, Rui Men, Hao Yang +15
Generalist models, which are capable of performing diverse multi-modal tasks in a task-agnostic way within a single model, have been explored recently. Being, hopefully, an alterna…
MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition
Xiaohuan Zhou, Jiaming Wang, Zeyu Cui +4
In this paper, we propose a novel multi-modal multi-task encoder-decoder pre-training framework (MMSpeech) for Mandarin automatic speech recognition (ASR), which employs both unlab…
Dimensionality-Varying Diffusion Process
Han Zhang, Ruili Feng, Zhantao Yang +7
Diffusion models, which learn to reverse a signal destruction process to generate new data, typically require the signal at each step to have the same dimension. We argue that, con…
Neural Dependencies Emerging from Learning Massive Categories
Ruili Feng, Kecheng Zheng, Kai Zhu +7
This work presents two astonishing findings on neural networks learned for large-scale image classification. 1) Given a well-trained model, the logits predicted for some category c…
Principled Knowledge Extrapolation with GANs
Ruili Feng, Jie Xiao, Kecheng Zheng +4
Human can extrapolate well, generalize daily knowledge into unseen scenarios, raise and answer counterfactual questions. To imitate this ability via generative models, previous wor…
M6-Rec: Generative Pretrained Language Models are Open-Ended Recommender Systems
Zeyu Cui, Jianxin Ma, Chang Zhou +2
Industrial recommender systems have been growing increasingly complex, may involve \emph{diverse domains} such as e-commerce products and user-generated contents, and can comprise…