20 citations · 63 across the 13 of their papers we have counts for
5 papers · 1 filter
On the Difference of BERT-style and CLIP-style Text Encoders
Zhihong Chen, Guiming Hardy Chen, Shizhe Diao +2
Masked language modeling (MLM) has been one of the most popular pretraining recipes in natural language processing, e.g., BERT, one of the representative models. Recently, contrast…
HuatuoGPT, towards Taming Language Model to Be a Doctor
Hongbo Zhang, Junying Chen, Feng Jiang +10
In this paper, we present HuatuoGPT, a large language model (LLM) for medical consultation. The core recipe of HuatuoGPT is to leverage both \textit{distilled data from ChatGPT} an…
Injecting Knowledge into Biomedical Pre-trained Models via Polymorphism and Synonymous Substitution
Hongbo Zhang, Xiang Wan, Benyou Wang
Pre-trained language models (PLMs) were considered to be able to store relational knowledge present in the training data. However, some relational knowledge seems to be discarded u…
Huatuo-26M, a Large-scale Chinese Medical QA Dataset
Jianquan Li, Xidong Wang, Xiangbo Wu +6
In this paper, we release a largest ever medical Question Answering (QA) dataset with 26 million QA pairs. We benchmark many existing approaches in our dataset in terms of both ret…
Phoenix: Democratizing ChatGPT across Languages
Zhihong Chen, Feng Jiang, Junying Chen +11
This paper presents our efforts to democratize ChatGPT across language. We release a large language model "Phoenix", achieving competitive performance among open-source English and…