papers

Publications (9)

cs.LG2024

Precise Model Benchmarking with Only a Few Observations

Riccardo Fogliato, Pratik Patil, Nil-Jana Akpinar +1

How can we precisely estimate a large language model's (LLM) accuracy on questions belonging to a specific topic within a larger question-answering dataset? The standard direct est…

cs.CV2019

Moments in Time Dataset: one million videos for event understanding

Mathew Monfort, Alex Andonian, Bolei Zhou +8

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modelin…

cs.CV2019

Reasoning About Human-Object Interactions Through Dual Attention Networks

Tete Xiao, Quanfu Fan, Dan Gutfreund +3

Objects are entities we act upon, where the functionality of an object is determined by how we interact with it. In this work we propose a Dual Attention Network model which reason…

cs.CV2019

Multi-Agent Tensor Fusion for Contextual Trajectory Prediction

Tianyang Zhao, Yifei Xu, Mathew Monfort +5

Accurate prediction of others' trajectories is essential for autonomous driving. Trajectory prediction is challenging because it requires reasoning about agents' past movements, so…

cs.CV2020

We Have So Much In Common: Modeling Semantic Relational Set Abstractions in Videos

Alex Andonian, Camilo Fosco, Mathew Monfort +4

Identifying common patterns among events is a key ability in human and machine perception, as it underlies intelligent decision making. We propose an approach for learning semantic…

cs.CV2021

Spoken Moments: Learning Joint Audio-Visual Representations from Video Descriptions

Mathew Monfort, SouYoung Jin, Alexander Liu +4

When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information…