2 citations · 7 across the 11 of their papers we have counts for
17 papers
QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling
Michael Y. Li, Anthony Zhan, Kanishk Gandhi +2
Scaling inference compute, by generating many parallel attempts per problem, is a costly but reliable lever for improving language model capabilities. By default these attempts are…
Learning Next Action Predictors from Human-Computer Interaction
Omar Shaikh, Valentin Teutschbein, Kanishk Gandhi +8
Truly proactive AI systems must anticipate what we will do next. This foresight demands far richer information than the sparse signals we type into our prompts -- it demands reason…
Endless Terminals: Scaling RL Environments for Terminal Agents
Kanishk Gandhi, Shivam Garg, Noah D. Goodman +1
Environments are the bottleneck for self-improving agents. Current terminal benchmarks were built for evaluation, not training; reinforcement learning requires a scalable pipeline,…
Learning to Simulate Human Dialogue
Kanishk Gandhi, Agam Bhatia, Noah D. Goodman
To predict what someone will say is to model how they think. We study this through next-turn dialogue prediction: given a conversation, predict the next utterance produced by a per…
Scaling up the think-aloud method
Daniel Wurgaft, Ben Prystawski, Kanishk Gandhi +3
The think-aloud method, where participants voice their thoughts as they solve a task, is a valuable source of rich data about human reasoning processes. Yet, it has declined in pop…
Non-literal Understanding of Number Words by Language Models
Polina Tsvilodub, Kanishk Gandhi, Haoran Zhao +3
Humans naturally interpret numbers non-literally, effortlessly combining context, world knowledge, and speaker intent. We investigate whether large language models (LLMs) interpret…