activity
20182024
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 1.4k across the 29 of their papers we have counts for

collaborators
Showing 2023 · cs.CLShow all

5 papers · 2 filters

cs.CL2023

BLESS: Benchmarking Large Language Models on Sentence Simplification

Tannon Kew, Alison Chi, Laura Vásquez-Rodríguez +4

We present BLESS, a comprehensive performance benchmark of the most recent state-of-the-art large language models (LLMs) on the task of text simplification (TS). We examine how wel…

cs.CL2023★ 5 cited

Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?

Xiangru Tang, Yiming Zong, Jason Phang +4

Despite the remarkable capabilities of Large Language Models (LLMs) like GPT-4, producing complex, structured tabular data remains challenging. Our study assesses LLMs' proficiency…

cs.CL2023★ 7 cited

Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs

Angelica Chen, Jason Phang, Alicia Parrish +4

Large language models (LLMs) have achieved widespread success on a variety of in-context few-shot tasks, but this success is typically evaluated via correctness rather than consist…

cs.CL2023★ 33 cited

Tool Learning with Foundation Models

Yujia Qin, Shengding Hu, Yankai Lin +38

Humans possess an extraordinary ability to create and utilize tools, allowing them to overcome physical limitations and explore new frontiers. With the advent of foundation models,…

cs.CL2023★ 26 cited

Pretraining Language Models with Human Preferences

Tomasz Korbak, Kejian Shi, Angelica Chen +5

Language models (LMs) are pretrained to imitate internet text, including content that would violate human preferences if generated by an LM: falsehoods, offensive comments, persona…