activity
20202023
most citedPre-Training a Language Model Without Human Language

3 citations · 7 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2023

Revealing the Blind Spot of Sentence Encoder Evaluation by HEROS

Cheng-Han Chiang, Yung-Sung Chuang, James Glass +1

Existing sentence textual similarity benchmark datasets only use a single number to summarize how similar the sentence encoder's decision is to humans'. However, it is unclear what…

cs.CL20223 cited

Re-Examining Human Annotations for Interpretable NLP

Cheng-Han Chiang, Hung-yi Lee

Explanation methods in Interpretable NLP often explain the model's decision by extracting evidence (rationale) from the input texts supporting the decision. Benchmark datasets for…

cs.CL20221 cited

Understanding, Detecting, and Separating Out-of-Distribution Samples and Adversarial Samples in Text Classification

Cheng-Han Chiang, Hung-yi Lee

In this paper, we study the differences and commonalities between statistically out-of-distribution (OOD) samples and adversarial (Adv) samples, both of which hurting a text classi…

cs.CL20203 cited

Pre-Training a Language Model Without Human Language

Cheng-Han Chiang, Hung-yi Lee

In this paper, we study how the intrinsic nature of pre-training data contributes to the fine-tuned downstream performance. To this end, we pre-train different transformer-based ma…

cs.CL2020

Pretrained Language Model Embryology: The Birth of ALBERT

Cheng-Han Chiang, Sung-Feng Huang, Hung-yi Lee

While behaviors of pretrained language models (LMs) have been thoroughly examined, what happened during pretraining is rarely studied. We thus investigate the developmental process…