17 citations · 37 across the 13 of their papers we have counts for
17 papers · 1 filter
What Makes Data-to-Text Generation Hard for Pretrained Language Models?
Moniba Keymanesh, Adrian Benton, Mark Dredze
Expressing natural language descriptions of structured facts or relations -- data-to-text generation (D2T) -- increases the accessibility of structured knowledge repositories. Prev…
Enriching Unsupervised User Embedding via Medical Concepts
Xiaolei Huang, Franck Dernoncourt, Mark Dredze
Clinical notes in Electronic Health Records (EHR) present rich documented information of patients to inference phenotype for disease diagnosis and study patient characteristics for…
Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information Extraction
Mahsa Yarmohammadi, Shijie Wu, Marc Marone +10
Zero-shot cross-lingual information extraction (IE) describes the construction of an IE model for some target language, given existing annotations exclusively in some other languag…
Learning to Look Inside: Augmenting Token-Based Encoders with Character-Level Information
Yuval Pinter, Amanda Stent, Mark Dredze +1
Commonly-used transformer language models depend on a tokenization schema which sets an unchangeable subword vocabulary prior to pre-training, destined to be applied to all downstr…
User Factor Adaptation for User Embedding via Multitask Learning
Xiaolei Huang, Michael J. Paul, Robin Burke +2
Language varies across users and their interested fields in social media data: words authored by a user across his/her interests may have different meanings (e.g., cool) or sentime…
Generating Synthetic Text Data to Evaluate Causal Inference Methods
Zach Wood-Doughty, Ilya Shpitser, Mark Dredze
Drawing causal conclusions from observational data requires making assumptions about the true data-generating process. Causal inference research typically considers low-dimensional…