activity
20172021
most citedSemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation

596 citations · 698 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

15 papers · 1 filter

cs.CL2021

A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations

Ziyi Yang, Yinfei Yang, Daniel Cer +1

Language agnostic and semantic-language information isolation is an emerging research direction for multilingual representations models. We explore this problem from a novel angle…

cs.CL2021

NT5?! Training T5 to Perform Numerical Reasoning

Peng-Jian Yang, Ying Ting Chen, Yuechan Chen +1

Numerical reasoning over text (NRoT) presents unique challenges that are not well addressed by existing pre-training objectives. We explore five sequential training schedules that…

cs.CL2020

Universal Sentence Representation Learning with Conditional Masked Language Model

Ziyi Yang, Yinfei Yang, Daniel Cer +2

This paper presents a novel training method, Conditional Masked Language Modeling (CMLM), to effectively learn sentence representations on large scale unlabeled corpora. CMLM integ…

cs.CL20202 cited

Neural Retrieval for Question Answering with Cross-Attention Supervised Data Augmentation

Yinfei Yang, Ning Jin, Kuo Lin +2

Neural models that independently project questions and answers into a shared embedding space allow for efficient continuous space retrieval from large corpora. Independently comput…

cs.CL202015 cited

MultiReQA: A Cross-Domain Evaluation for Retrieval Question Answering Models

Mandy Guo, Yinfei Yang, Daniel Cer +2

Retrieval question answering (ReQA) is the task of retrieving a sentence-level answer to a question from an open corpus (Ahmad et al.,2019).This paper presents MultiReQA, anew mult…

cs.CL2020

Crisscrossed Captions: Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCO

Zarana Parekh, Jason Baldridge, Daniel Cer +2

By supporting multi-modal retrieval training and evaluation, image captioning datasets have spurred remarkable progress on representation learning. Unfortunately, datasets have lim…