most citedWhen Do You Need Billions of Words of Pretraining Data?

16 citations · 24 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL202016 cited

When Do You Need Billions of Words of Pretraining Data?

Yian Zhang, Alex Warstadt, Haau-Sing Li +1

NLP is currently dominated by general-purpose pretrained language models like RoBERTa, which achieve strong performance on NLU tasks through pretraining on billions of words. But w…

cs.CL2020

Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)

Alex Warstadt, Yian Zhang, Haau-Sing Li +2

One reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding. However, we want pretrained…

cs.CL2020

Latent Tree Learning with Ordered Neurons: What Parses Does It Produce?

Yian Zhang

Recent latent tree learning models can learn constituency parsing without any exposure to human-annotated tree structures. One such model is ON-LSTM (Shen et al., 2019), which is t…

cs.HC20201 cited

Interactive Rainbow Score: A Visual-centered Multimodal Flute Tutoring System

Daniel Chin, Yian Zhang, Tianyu Zhang +2

Learning to play an instrument is intrinsically multimodal, and we have seen a trend of applying visual and haptic feedback in music games and computer-aided music tutoring systems…

cs.HC20197 cited

Adaptive Multimodal Music Learning via Interactive-haptic Instrument

Yian Zhang, Yinmiao Li, Daniel Chin +1

Haptic interfaces have untapped the sense of touch to assist multimodal music learning. We have recently seen various improvements of interface design on tactile feedback and force…