activity
20102022
most citedEfficiently Inducing Features of Conditional Random Fields

366 citations · 950 across the 41 of their papers we have counts for

collaborators
Showing cs.CLShow all

45 papers · 1 filter

cs.CL20223 cited

You can't pick your neighbors, or can you? When and how to rely on retrieval in the NN-LM

Andrew Drozdov, Shufan Wang, Razieh Rahimi +3

Retrieval-enhanced language models (LMs), which condition their predictions on text retrieved from large external datastores, have recently shown significant perplexity improvement…

cs.CL2022

Efficient Nearest Neighbor Search for Cross-Encoder Models using Matrix Factorization

Nishant Yadav, Nicholas Monath, Rico Angell +2

Efficient k-nearest neighbor search is a fundamental task, foundational for many problems in NLP. When the similarity is measured by dot-product between dual-encoder vectors or $\e…

cs.CL2022

Longtonotes: OntoNotes with Longer Coreference Chains

Kumar Shridhar, Nicholas Monath, Raghuveer Thirukovalluru +4

Ontonotes has served as the most important benchmark for coreference resolution. However, for ease of annotation, several long documents in Ontonotes were split into smaller parts.…

cs.CL20221 cited

Inducing and Using Alignments for Transition-based AMR Parsing

Andrew Drozdov, Jiawei Zhou, Radu Florian +4

Transition-based parsers for Abstract Meaning Representation (AMR) rely on node-to-word alignments. These alignments are learned separately from parser training and require a compl…

cs.CL20228 cited

CBR-iKB: A Case-Based Reasoning Approach for Question Answering over Incomplete Knowledge Bases

Dung Thai, Srinivas Ravishankar, Ibrahim Abdelaziz +7

Knowledge bases (KBs) are often incomplete and constantly changing in practice. Yet, in many question answering applications coupled with knowledge bases, the sparse nature of KBs…

cs.CL2022

A Distant Supervision Corpus for Extracting Biomedical Relationships Between Chemicals, Diseases and Genes

Dongxu Zhang, Sunil Mohan, Michaela Torkar +1

We introduce ChemDisGene, a new dataset for training and evaluating multi-class multi-label document-level biomedical relation extraction models. Our dataset contains 80k biomedica…