activity
20212023
most citedLook at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

13 citations · 14 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI20231 cited

Socratis: Are large multimodal models emotionally aware?

Katherine Deng, Arijit Ray, Reuben Tan +3

Existing emotion prediction benchmarks contain coarse emotion labels which do not consider the diversity of emotions that an image and text can elicit in humans due to various reas…

cs.CV2023

Multiscale Video Pretraining for Long-Term Activity Forecasting

Reuben Tan, Matthias De Lange, Michael Iuzzolino +4

Long-term activity forecasting is an especially challenging research problem because it requires understanding the temporal relationships between observed actions, as well as the v…

cs.CV2023

EgoAdapt: A multi-stream evaluation study of adaptation to real-world egocentric user video

Matthias De Lange, Hamid Eghbalzadeh, Reuben Tan +3

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this…

cs.CV2022

NewsStories: Illustrating articles with visual summaries

Reuben Tan, Bryan A. Plummer, Kate Saenko +3

Recent self-supervised approaches have used large-scale image-text datasets to learn powerful representations that transfer to many tasks without finetuning. These methods often as…

cs.CV202113 cited

Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

Reuben Tan, Bryan A. Plummer, Kate Saenko +2

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision…