activity
20172026
most citedRevisiting Distributed Synchronous SGD

609 citations · 1.3k across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection

Atoosa Chegini, Hamid Kazemi, Garrett Souza +5

Reasoning has become a central paradigm for large language models (LLMs), consistently boosting accuracy across diverse benchmarks. Yet its suitability for precision-sensitive task…

cs.CL2025

AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking

Silin Gao, Antoine Bosselut, Samy Bengio +1

Recent studies have shown that large language models (LLMs), especially smaller ones, often lack robustness in grade school math (GSM) reasoning. In particular, they tend to experi…

cs.CL2025

What Makes the Preferred Thinking Direction for LLMs in Multiple-choice Questions?

Yizhe Zhang, Richard Bai, Zijin Gu +5

Language models usually use left-to-right (L2R) autoregressive factorization. However, L2R factorization may not always be the best inductive bias. Therefore, we investigate whethe…

cs.CL2019

Parallel Scheduled Sampling

Daniel Duckworth, Arvind Neelakantan, Ben Goodrich +2

Auto-regressive models are widely used in sequence generation problems. The output sequence is typically generated in a predetermined order, one discrete unit (pixel or word or cha…

cs.CL2018

Content preserving text generation with attribute controls

Lajanugen Logeswaran, Honglak Lee, Samy Bengio

In this work, we address the problem of modifying textual attributes of sentences. Given an input sentence and a set of attribute labels, we attempt to generate sentences that are…