activity
20172024
most citedCyclical Annealing Schedule: A Simple Approach to Mitigating KL Vanishing

169 citations · 738 across the 46 of their papers we have counts for

collaborators
Showing cs.CLShow all

47 papers · 1 filter

cs.CL2024★ 5 cited

Efficient Tool Use with Chain-of-Abstraction Reasoning

Silin Gao, Jane Dwivedi-Yu, Ping Yu +7

To achieve faithful reasoning that aligns with human expectations, large language models (LLMs) need to ground their reasoning to real-world knowledge (e.g., web facts, math and ph…

cs.CL2023

PathFinder: Guided Search over Multi-Step Reasoning Paths

Olga Golovneva, Sean O'Brien, Ramakanth Pasunuru +4

With recent advancements in large language models, methods like chain-of-thought prompting to elicit reasoning chains have been shown to improve results on reasoning tasks. However…

cs.CL2023★ 1 cited

The ART of LLM Refinement: Ask, Refine, and Trust

Kumar Shridhar, Koustuv Sinha, Andrew Cohen +6

In recent years, Large Language Models (LLMs) have demonstrated remarkable generative abilities, but can they judge the quality of their own generations? A popular concept, referre…

cs.CL2023

BLESS: Benchmarking Large Language Models on Sentence Simplification

Tannon Kew, Alison Chi, Laura Vásquez-Rodríguez +4

We present BLESS, a comprehensive performance benchmark of the most recent state-of-the-art large language models (LLMs) on the task of text simplification (TS). We examine how wel…

cs.CL2023★ 1 cited

Sub-network Discovery and Soft-masking for Continual Learning of Mixed Tasks

Zixuan Ke, Bing Liu, Wenhan Xiong +2

Continual learning (CL) has two main objectives: preventing catastrophic forgetting (CF) and encouraging knowledge transfer (KT). The existing literature mainly focused on overcomi…

cs.CL2023★ 9 cited

Walking Down the Memory Maze: Beyond Context Limit through Interactive Reading

Howard Chen, Ramakanth Pasunuru, Jason Weston +1

Large language models (LLMs) have advanced in large strides due to the effectiveness of the self-attention mechanism that processes and compares all tokens at once. However, this m…