activity
20142022
most citedBucking the Trend: Large-Scale Cost-Focused Active Learning for Statistical Machine Translation

62 citations · 122 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL20152 cited

Annotating Cognates and Etymological Origin in Turkic Languages

Benjamin S. Mericli, Michael Bloodgood

Turkic languages exhibit extensive and diverse etymological relationships among lexical items. These relationships make the Turkic languages promising for exploring automated trans…

cs.CL2014

Rapid Adaptation of POS Tagging for Domain Specific Uses

John E. Miller, Michael Bloodgood, Manabu Torii +1

Part-of-speech (POS) tagging is a fundamental component for performing natural language tasks such as parsing, information extraction, and question answering. When POS taggers are…

cs.CL20142 cited

A random forest system combination approach for error detection in digital dictionaries

Michael Bloodgood, Peng Ye, Paul Rodrigues +2

When digitizing a print bilingual dictionary, whether via optical character recognition or manual entry, it is inevitable that errors are introduced into the electronic version tha…

cs.CL20141 cited

Correcting Errors in Digital Lexicographic Resources Using a Dictionary Manipulation Language

David Zajic, Michael Maxwell, David Doermann +2

We describe a paradigm for combining manual and automatic error correction of noisy structured lexicographic data. Modifications to the structure and underlying text of the lexicog…

cs.CL201462 cited

Bucking the Trend: Large-Scale Cost-Focused Active Learning for Statistical Machine Translation

Michael Bloodgood, Chris Callison-Burch

We explore how to improve machine translation systems by adding more translation data in situations where we already have substantial resources. The main challenge is how to buck t…

cs.CL201435 cited

Using Mechanical Turk to Build Machine Translation Evaluation Sets

Michael Bloodgood, Chris Callison-Burch

Building machine translation (MT) test sets is a relatively expensive task. As MT becomes increasingly desired for more and more language pairs and more and more domains, it become…