4 citations · 4 across the 4 of their papers we have counts for
7 papers · 1 filter
Learning Model Successors
Yingshan Chang, Yonatan Bisk
The notion of generalization has moved away from the classical one defined in statistical learning theory towards an emphasis on out-of-domain generalization (OODG). There has been…
Looking beyond the next token
Abitha Thankaraj, Yiding Jiang, J. Zico Kolter +1
The structure of causal language model training assumes that each token can be accurately predicted from the previous context. This contrasts with humans' natural writing and reaso…
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
Jared Fernandez, Luca Wehrstedt, Leonid Shamis +5
Dramatic increases in the capabilities of neural network models in recent years are driven by scaling model size, training data, and corresponding computational resources. To devel…
Self-Regulation and Requesting Interventions
So Yeon Min, Yue Wu, Jimin Sun +4
Human intelligence involves metacognitive abilities like self-regulation, recognizing limitations, and seeking assistance only when needed. While LLM Agents excel in many domains,…
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
Melanie Sclar, Jane Yu, Maryam Fazel-Zarandi +4
Do large language models (LLMs) have theory of mind? A plethora of papers and benchmarks have been introduced to evaluate if current models have been able to develop this key abili…
Language Models Need Inductive Biases to Count Inductively
Yingshan Chang, Yonatan Bisk
Counting is a fundamental example of generalization, whether viewed through the mathematical lens of Peano's axioms defining the natural numbers or the cognitive science literature…