papers

Publications (30)

cs.CL2022

BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric

Mingda Chen, Paul-Ambroise Duquenne, Pierre Andrews +4

End-to-End speech-to-speech translation (S2ST) is generally evaluated with text-based metrics. This means that generated speech has to be automatically transcribed, making the eval…

cs.CL2020

Learning Probabilistic Sentence Representations from Paraphrases

Mingda Chen, Kevin Gimpel

Probabilistic word embeddings have shown effectiveness in capturing notions of generality and entailment, but there is very little work on doing the analogous type of investigation…

cs.CL2024

Few-Shot Data Synthesis for Open Domain Multi-Hop Question Answering

Mingda Chen, Xilun Chen, Wen-tau Yih

Few-shot learning for open domain multi-hop question answering typically relies on the incontext learning capability of large language models (LLMs). While powerful, these LLMs usu…

cs.CL2019

A Multi-Task Approach for Disentangling Syntax and Semantics in Sentence Representations

Mingda Chen, Qingming Tang, Sam Wiseman +1

We propose a generative model for a sentence that uses two latent variables, with one intended to represent the syntax of the sentence and the other to represent its semantics. We…

cs.CL2026

A Survey on Diffusion Language Models

Tianyi Li, Mingda Chen, Bowei Guo +1

Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm. By generating tokens in parallel through…

cs.CL2019

How to Ask Better Questions? A Large-Scale Multi-Domain Dataset for Rewriting Ill-Formed Questions

Zewei Chu, Mingda Chen, Jing Chen +4

We present a large-scale dataset for the task of rewriting an ill-formed natural language question to a well-formed one. Our multi-domain question rewriting MQR dataset is construc…