32 citations · 99 across the 15 of their papers we have counts for
14 papers · 1 filter
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Shunyu Yao, Dian Yu, Jeffrey Zhao +4
Language models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making proce…
Can Rationalization Improve Robustness?
Howard Chen, Jacqueline He, Karthik Narasimhan +1
A growing line of work has investigated the development of neural NLP models that can produce rationales--subsets of input that can explain their model predictions. In this paper,…
CARETS: A Consistency And Robustness Evaluative Test Suite for VQA
Carlos E. Jimenez, Olga Russakovsky, Karthik Narasimhan
We introduce CARETS, a systematic test suite to measure consistency and robustness of modern VQA models through a series of six fine-grained capability tests. In contrast to existi…
Reading and Acting while Blindfolded: The Need for Semantics in Text Game Agents
Shunyu Yao, Karthik Narasimhan, Matthew Hausknecht
Text-based games simulate worlds and interact with players using natural language. Recent work has used them as a testbed for autonomous language-understanding agents, with the mot…
Grounding Language to Entities and Dynamics for Generalization in Reinforcement Learning
Austin W. Hanjie, Victor Zhong, Karthik Narasimhan
We investigate the use of natural language to drive the generalization of control policies and introduce the new multi-task environment Messenger with free-form text manuals descri…
Keep CALM and Explore: Language Models for Action Generation in Text-based Games
Shunyu Yao, Rohan Rao, Matthew Hausknecht +1
Text-based games present a unique challenge for autonomous agents to operate in natural language and handle enormous action spaces. In this paper, we propose the Contextual Action…