Publications (52)
Synthetic QA Corpora Generation with Roundtrip Consistency
Chris Alberti, Daniel Andor, Emily Pitler +2
Superflows: A New Tool for Forensic Network Flow Analysis
Michael Collins, Jyotirmoy V. Deshmukh, Dristi Dinesh +3
Sparse, Dense, and Attentional Representations for Text Retrieval
Yi Luan, Jacob Eisenstein, Kristina Toutanova +1
Learning Dictionaries for Named Entity Recognition using Minimal Supervision
Arvind Neelakantan, Michael Collins
The Lowell Observatory Solar Telescope: A fiber feed into the EXtreme PREcision Spectrometer
Joe Llama, Lily L. Zhao, John M. Brewer +5
Controlled Decoding from Language Models
Sidharth Mudgal, Jong Lee, Harish Ganapathy +10
Evaluating Explanations: How much do explanations from the teacher aid students?
Danish Pruthi, Rachit Bansal, Bhuwan Dhingra +5
Partially Supervised Named Entity Recognition via the Expected Entity Ratio Loss
Thomas Effland, Michael Collins
QED: A Framework and Dataset for Explanations in Question Answering
Matthew Lamm, Jennimaria Palomaki, Chris Alberti +4
Kernel Approximation Methods for Speech Recognition
Avner May, Alireza Bagheri Garakani, Zhiyun Lu +9
Failure of Optimal Design Theory? A Case Study in Toxicology Using Sequential Robust Optimal Design Framework
Elvis Han Cui, Michael Collins, Jessica Munson +1
Case-Factor Diagrams for Structured Probabilistic Modeling
David A. McAllester, Michael Collins, Fernando Pereira
Decontextualization: Making Sentences Stand-Alone
Eunsol Choi, Jennimaria Palomaki, Matthew Lamm +3
Query Refinement Prompts for Closed-Book Long-Form Question Answering
Reinald Kim Amplayo, Kellie Webster, Michael Collins +2
Fusion of Detected Objects in Text for Visual Question Answering
Chris Alberti, Jeffrey Ling, Michael Collins +1
On planetary systems as ordered sequences
Emily Sandford, David Kipping, Michael Collins
A Sociotechnical, Practitioner-Centered Approach to Technology Adoption in Cybersecurity Operations: An LLM Case
Francis Hahn, Mohd Mamoon, Alexandru G. Bardas +4
How to Scale Up Kernel Methods to Be As Good As Deep Neural Nets
Zhiyun Lu, Avner May, Kuan Liu +8
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
Bernd Bohnet, Kevin Swersky, Rosanne Liu +9
Honest Students from Untrusted Teachers: Learning an Interpretable Question-Answering Pipeline from a Pretrained Language Model
Jacob Eisenstein, Daniel Andor, Bernd Bohnet +2
The multiplicity distribution of Kepler's exoplanets
Emily Sandford, David Kipping, Michael Collins
Prepositional Phrase Attachment through a Backed-Off Model
Michael Collins, James Brooks
Improving Low-Resource Cross-lingual Parsing with Expected Statistic Regularization
Thomas Effland, Michael Collins
Globally Normalized Transition-Based Neural Networks
Daniel Andor, Chris Alberti, David Weiss +5
Learning to Reject with a Fixed Predictor: Application to Decontextualization
Christopher Mohri, Daniel Andor, Eunsol Choi +1
Existence of parabolic minimizers to the total variation flow on metric measure spaces
Vito Buffa, Michael Collins, Cintia Pacchiano Camacho
Structured Training for Neural Network Transition-Based Parsing
David Weiss, Chris Alberti, Michael Collins +1
Cross-Lingual Syntactic Transfer with Limited Resources
Mohammad Sadegh Rasooli, Michael Collins
MEG-Derived Functional Tractography, Results for Normal and Concussed Cohorts
Don Krieger, Paul Shepard, Walter Schneider +4
A New Statistical Parser Based on Bigram Lexical Dependencies
Michael Collins
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
Towards Computationally Verifiable Semantic Grounding for Language Models
Chris Alberti, Kuzman Ganchev, Michael Collins +2
Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
Bernd Bohnet, Vinh Q. Tran, Pat Verga +19
A Well-Composed Text is Half Done! Composition Sampling for Diverse Conditional Generation
Shashi Narayan, Gonçalo Simões, Yao Zhao +4
Low-Resource Syntactic Transfer with Unsupervised Source Reordering
Mohammad Sadegh Rasooli, Michael Collins
Three Generative, Lexicalised Models for Statistical Parsing
Michael Collins
Coreference Resolution through a seq2seq Transition-Based System
Bernd Bohnet, Chris Alberti, Michael Collins
Improving Span-based Question Answering Systems with Coarsely Labeled Data
Hao Cheng, Ming-Wei Chang, Kenton Lee +3
A Tutorial on Dual Decomposition and Lagrangian Relaxation for Inference in Natural Language Processing
Alexander M. Rush, Michael Collins
Learning to Map Sentences to Logical Form: Structured Classification with Probabilistic Categorial Grammars
Luke S. Zettlemoyer, Michael Collins
A Comparison between Deep Neural Nets and Kernel Acoustic Models for Speech Recognition
Zhiyun Lu, Dong Guo, Alireza Bagheri Garakani +8
Hunting in the Dark: Metrics for Early Stage Traffic Discovery
Max Gao, Michael Collins, Ricky Mok +1
NeurIPS 2020 EfficientQA Competition: Systems, Analyses and Lessons Learned
Sewon Min, Jordan Boyd-Graber, Chris Alberti +50
InfAlign: Inference-aware language model alignment
Ananth Balashankar, Ziteng Sun, Jonathan Berant +9
SyntaxNet Models for the CoNLL 2017 Shared Task
Chris Alberti, Daniel Andor, Ivan Bogatyy +10
A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
Alon Jacovi, Yonatan Bitton, Bernd Bohnet +6
High-Fidelity Ion State Detection Using Trap-Integrated Avalanche Photodiodes
David Reens, Michael Collins, Joseph Ciampi +12
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
Jonathan H. Clark, Eunsol Choi, Michael Collins +4
A BERT Baseline for the Natural Questions
Chris Alberti, Kenton Lee, Michael Collins
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Christopher Clark, Kenton Lee, Ming-Wei Chang +3
Noise Contrastive Estimation and Negative Sampling for Conditional Models: Consistency and Statistical Efficiency
Zhuang Ma, Michael Collins
Measuring Attribution in Natural Language Generation Models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm +7