activity
20162026
most citedNo Training Required: Exploring Random Encoders for Sentence Classification

75 citations · 135 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

16 papers · 1 filter

cs.CL20263 cited

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…

cs.CL2026

StoryScope: Investigating idiosyncrasies in AI fiction

Jenna Russell, Rishanth Rajendhran, Chau Minh Pham +2

As AI-generated fiction becomes increasingly prevalent, questions of authorship and originality are becoming central to how written work is evaluated. While most existing work in t…

cs.CL2024

Learning from Many Voices: Literary MT Using Multi-Reference Human and Synthetic Data

Si Wu, John Wieting, David A. Smith

Multiple valid translations of a single literary work naturally exist. We investigate strategies for leveraging these multi-reference datasets to improve literary machine translati…

cs.CL20223 cited

Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature

Katherine Thai, Marzena Karpinska, Kalpesh Krishna +4

Literary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators relative to the many untranslated works publishe…

cs.CL20221 cited

Faithful to the Document or to the World? Mitigating Hallucinations via Entity-linked Knowledge in Abstractive Summarization

Yue Dong, John Wieting, Pat Verga

Despite recent advances in abstractive summarization, current summarization systems still suffer from content hallucinations where models generate text that is either irrelevant or…

cs.CL2021

Improving the Diversity of Unsupervised Paraphrasing with Embedding Outputs

Monisha Jegadeesan, Sachin Kumar, John Wieting +1

We present a novel technique for zero-shot paraphrase generation. The key contribution is an end-to-end multilingual paraphrasing model that is trained using translated parallel co…