activity
20162021
collaborators

5 papers

cs.CV2021

Spoken Moments: Learning Joint Audio-Visual Representations from Video Descriptions

Mathew Monfort, SouYoung Jin, Alexander Liu +4

When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information…

cs.CV2020

We Have So Much In Common: Modeling Semantic Relational Set Abstractions in Videos

Alex Andonian, Camilo Fosco, Mathew Monfort +4

Identifying common patterns among events is a key ability in human and machine perception, as it underlies intelligent decision making. We propose an approach for learning semantic…

cs.CV2019

Reasoning About Human-Object Interactions Through Dual Attention Networks

Tete Xiao, Quanfu Fan, Dan Gutfreund +3

Objects are entities we act upon, where the functionality of an object is determined by how we interact with it. In this work we propose a Dual Attention Network model which reason…

cs.CV2019

Multi-Agent Tensor Fusion for Contextual Trajectory Prediction

Tianyang Zhao, Yifei Xu, Mathew Monfort +5

Accurate prediction of others' trajectories is essential for autonomous driving. Trajectory prediction is challenging because it requires reasoning about agents' past movements, so…

cs.CV2016

End to End Learning for Self-Driving Cars

Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski +10

We trained a convolutional neural network (CNN) to map raw pixels from a single front-facing camera directly to steering commands. This end-to-end approach proved surprisingly powe…