most citedBigBIO: A Framework for Data-Centric Biomedical Natural Language Processing

12 citations · 22 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL20231 cited

Multi-lingual and Multi-cultural Figurative Language Understanding

Anubha Kabra, Emmy Liu, Simran Khanuja +6

Figurative language permeates human communication, but at the same time is relatively understudied in NLP. Datasets have been created in English to accelerate progress towards meas…

cs.CL20232 cited

GlobalBench: A Benchmark for Global Progress in Natural Language Processing

Yueqi Song, Catherine Cui, Simran Khanuja +9

Despite the major advances in NLP, significant disparities in NLP system performance across languages still exist. Arguably, these are due to uneven resource allocation and sub-opt…

cs.CL2023

Which One Are You Referring To? Multimodal Object Identification in Situated Dialogue

Holy Lovenia, Samuel Cahyawijaya, Pascale Fung

The demand for multimodal dialogue systems has been rising in various domains, emphasizing the importance of interpreting multimodal inputs from conversational and situational cont…

cs.CL20224 cited

NusaCrowd: A Call for Open and Reproducible NLP Research in Indonesian Languages

Samuel Cahyawijaya, Alham Fikri Aji, Holy Lovenia +8

At the center of the underlying issues that halt Indonesian natural language processing (NLP) research advancement, we find data scarcity. Resources in Indonesian languages, especi…

cs.CL2022

Kaggle Competition: Cantonese Audio-Visual Speech Recognition for In-car Commands

Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu +2

With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-ca…

cs.CL202212 cited

BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing

Jason Alan Fries, Leon Weber, Natasha Seelam +40

Training and evaluating language models increasingly requires the construction of meta-datasets --diverse collections of curated data with clear provenance. Natural language prompt…