1.8k citations · 2.1k across the 20 of their papers we have counts for
6 papers · 1 filter
BRIDGE: Predicting Human Task Completion Time From Model Performance
Fengyuan Liu, Jay Gala, Nilaksh +3
Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on d…
TapeAgents: a Holistic Framework for Agent Development and Optimization
Dzmitry Bahdanau, Nicolas Gontier, Gabriel Huang +10
We present TapeAgents, an agent framework built around a granular, structured log tape of the agent session that also plays the role of the session's resumable state. In TapeAgents…
BabyAI 1.1
David Yu-Tung Hui, Maxime Chevalier-Boisvert, Dzmitry Bahdanau +1
The BabyAI platform is designed to measure the sample efficiency of training an agent to follow grounded-language instructions. BabyAI 1.0 presents baseline results of an agent tra…
CLOSURE: Assessing Systematic Generalization of CLEVR Models
Dzmitry Bahdanau, Harm de Vries, Timothy J. O'Donnell +4
The CLEVR dataset of natural-looking questions about 3D-rendered scenes has recently received much attention from the research community. A number of models have been proposed for…
BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou +4
Allowing humans to interactively train artificial agents to understand language instructions is desirable for both practical and scientific reasons, but given the poor data efficie…
Learning to Understand Goal Specifications by Modelling Reward
Dzmitry Bahdanau, Felix Hill, Jan Leike +4
Recent work has shown that deep reinforcement-learning agents can learn to follow language-like instructions from infrequent environment rewards. However, this places on environmen…