activity
20162021
collaborators

9 papers

cs.CL2021

Segmenting Subtitles for Correcting ASR Segmentation Errors

David Wan, Chris Kedzie, Faisal Ladhak +6

Typical ASR systems segment the input audio into utterances using purely acoustic information, which may not resemble the sentence-like units that are expected by conventional mach…

cs.CL2020

Subtitles to Segmentation: Improving Low-Resource Speech-to-Text Translation Pipelines

David Wan, Zhengping Jiang, Chris Kedzie +3

In this work, we focus on improving ASR output segmentation in the context of low-resource language speech-to-text translation. ASR output segmentation is crucial, as ASR systems s…

cs.CL2020

Incorporating Terminology Constraints in Automatic Post-Editing

David Wan, Chris Kedzie, Faisal Ladhak +2

Users of machine translation (MT) may want to ensure the use of specific lexical terminologies. While there exist techniques for incorporating terminology constraints during infere…

cs.CL2019

Low-Level Linguistic Controls for Style Transfer and Content Preservation

Katy Gero, Chris Kedzie, Jonathan Reeve +1

Despite the success of style transfer in image processing, it has seen limited progress in natural language generation. Part of the problem is that content is not as easily decoupl…

cs.CL2019

A Good Sample is Hard to Find: Noise Injection Sampling and Self-Training for Neural Language Generation Models

Chris Kedzie, Kathleen McKeown

Deep neural networks (DNN) are quickly becoming the de facto standard modeling method for many natural language generation (NLG) tasks. In order for such models to truly be useful,…

cs.CL2018

Content Selection in Deep Learning Models of Summarization

Chris Kedzie, Kathleen McKeown, Hal Daume

We carry out experiments with deep learning models of summarization across the domains of news, personal stories, meetings, and medical articles in order to understand how content…