activity
20162025
most citedMultilingual Universal Sentence Encoder for Semantic Retrieval

67 citations · 84 across the 5 of their papers we have counts for

collaborators

12 papers

cs.CL2025

Dr Genre: Reinforcement Learning from Decoupled LLM Feedback for Generic Text Rewriting

Yufei Li, John Nham, Ganesh Jawahar +7

Generic text rewriting is a prevalent large language model (LLM) application that covers diverse real-world tasks, such as style transfer, fact correction, and email editing. These…

cs.CL2023

Characterizing Tradeoffs in Language Model Decoding with Informational Interpretations

Chung-Ching Chang, William W. Cohen, Yun-Hsuan Sung

We propose a theoretical framework for formulating language model decoder algorithms with dynamic programming and information theory. With dynamic programming, we lift the design o…

cs.CL20223 cited

Knowledge Prompts: Injecting World Knowledge into Language Models through Soft Prompts

Cicero Nogueira dos Santos, Zhe Dong, Daniel Cer +4

Soft prompts have been recently proposed as a tool for adapting large frozen language models (LMs) to new tasks. In this work, we repurpose soft prompts to the task of injecting wo…

cs.CV2021

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Chao Jia, Yinfei Yang, Ye Xia +7

Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human anno…

cs.CL201967 cited

Multilingual Universal Sentence Encoder for Semantic Retrieval

Yinfei Yang, Daniel Cer, Amin Ahmad +9

We introduce two pre-trained retrieval focused multilingual sentence encoding models, respectively based on the Transformer and CNN model architectures. The models embed text from…

cs.CL20192 cited

Hierarchical Document Encoder for Parallel Corpus Mining

Mandy Guo, Yinfei Yang, Keith Stevens +5

We explore using multilingual document embeddings for nearest neighbor mining of parallel data. Three document-level representations are investigated: (i) document embeddings gener…