activity
20192022
most citedSame Object, Different Grasps: Data and Semantic Knowledge for Task-Oriented Grasping

18 citations · 24 across the 4 of their papers we have counts for

collaborators

6 papers

cs.LG20221 cited

Learning to Navigate Wikipedia by Taking Random Walks

Manzil Zaheer, Kenneth Marino, Will Grathwohl +7

A fundamental ability of an intelligent web-based agent is seeking out and acquiring new information. Internet search engines reliably find the correct vicinity but the top results…

cs.CV20205 cited

KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA

Kenneth Marino, Xinlei Chen, Devi Parikh +2

One of the most challenging question types in VQA is when answering the question requires outside knowledge not present in the image. In this work we study open-domain knowledge, t…

cs.RO202018 cited

Same Object, Different Grasps: Data and Semantic Knowledge for Task-Oriented Grasping

Adithyavairavan Murali, Weiyu Liu, Kenneth Marino +2

Despite the enormous progress and generalization in robotic grasping in recent years, existing methods have yet to scale and generalize task-oriented grasping to the same extent. T…

cs.LG2020

Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement Learning

Valerie Chen, Abhinav Gupta, Kenneth Marino

Complex, multi-task problems have proven to be difficult to solve efficiently in a sparse-reward reinforcement learning setting. In order to be sample efficient, multi-task learnin…

cs.AI2020

Empirically Verifying Hypotheses Using Reinforcement Learning

Kenneth Marino, Rob Fergus, Arthur Szlam +1

This paper formulates hypothesis verification as an RL problem. Specifically, we aim to build an agent that, given a hypothesis about the dynamics of the world, can take actions to…

cs.CV2019

OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Kenneth Marino, Mohammad Rastegari, Ali Farhadi +1

Visual Question Answering (VQA) in its ideal form lets us study reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. Ho…