activity
20142022
most citedEarly Recognition of Human Activities from First-Person Videos Using Onset Representations

8 citations · 22 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2024

Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed +2

Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly…

cs.CV2023

AAN: Attributes-Aware Network for Temporal Action Detection

Rui Dai, Srijan Das, Michael S. Ryoo +1

The challenge of long-term video understanding remains constrained by the efficient extraction of object semantics and the modelling of their relationships for downstream tasks. Al…

cs.CV2023

Energy-Based Models for Cross-Modal Localization using Convolutional Transformers

Alan Wu, Michael S. Ryoo

We present a novel framework using Energy-Based Models (EBMs) for localizing a ground vehicle mounted with a range sensor against satellite imagery in the absence of GPS. Lidar sen…

cs.CV20222 cited

Video Question Answering with Iterative Video-Text Co-Tokenization

AJ Piergiovanni, Kairo Morton, Weicheng Kuo +2

Video question answering is a challenging task that requires understanding jointly the language input, the visual information in individual video frames, as well as the temporal in…

cs.CV20227 cited

Video + CLIP Baseline for Ego4D Long-term Action Anticipation

Srijan Das, Michael S. Ryoo

In this report, we introduce our adaptation of image-text models for long-term action anticipation. Our Video + CLIP framework makes use of a large-scale pre-trained paired image-t…

cs.CV20212 cited

ViewCLR: Learning Self-supervised Video Representation for Unseen Viewpoints

Srijan Das, Michael S. Ryoo

Learning self-supervised video representation predominantly focuses on discriminating instances generated from simple data augmentation schemes. However, the learned representation…