8 citations · 8 across the 2 of their papers we have counts for
4 papers · 1 filter
What do Language Models Learn and When? The Implicit Curriculum Hypothesis
Emmy Liu, Kaiser Sun, Millicent Li +4
Large language models (LLMs) can perform remarkably complex tasks, yet the fine-grained details of how these capabilities emerge during pretraining remain poorly understood. Scalin…
On Retrieval Augmentation and the Limitations of Language Model Training
Ting-Rui Chiang, Xinyan Velocity Yu, Joshua Robinson +3
Augmenting a language model (LM) with -nearest neighbors (NN) retrieval on its training data alone can decrease its perplexity, though the underlying reasons for this remain…
Self-Contradictory Reasoning Evaluation and Detection
Ziyi Liu, Soumya Sanyal, Isabelle Lee +4
In a plethora of recent work, large language models (LLMs) demonstrated impressive reasoning ability, but many proposed downstream reasoning tasks only focus on final answers. Two…
GupShup: An Annotated Corpus for Abstractive Summarization of Open-Domain Code-Switched Conversations
Laiba Mehnaz, Debanjan Mahata, Rakesh Gosangi +7
Code-switching is the communication phenomenon where speakers switch between different languages during a conversation. With the widespread adoption of conversational agents and ch…