8 citations · 22 across the 6 of their papers we have counts for
9 papers · 1 filter
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed +2
Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly…
AAN: Attributes-Aware Network for Temporal Action Detection
Rui Dai, Srijan Das, Michael S. Ryoo +1
The challenge of long-term video understanding remains constrained by the efficient extraction of object semantics and the modelling of their relationships for downstream tasks. Al…
Energy-Based Models for Cross-Modal Localization using Convolutional Transformers
Alan Wu, Michael S. Ryoo
We present a novel framework using Energy-Based Models (EBMs) for localizing a ground vehicle mounted with a range sensor against satellite imagery in the absence of GPS. Lidar sen…
Video Question Answering with Iterative Video-Text Co-Tokenization
AJ Piergiovanni, Kairo Morton, Weicheng Kuo +2
Video question answering is a challenging task that requires understanding jointly the language input, the visual information in individual video frames, as well as the temporal in…
Video + CLIP Baseline for Ego4D Long-term Action Anticipation
Srijan Das, Michael S. Ryoo
In this report, we introduce our adaptation of image-text models for long-term action anticipation. Our Video + CLIP framework makes use of a large-scale pre-trained paired image-t…
ViewCLR: Learning Self-supervised Video Representation for Unseen Viewpoints
Srijan Das, Michael S. Ryoo
Learning self-supervised video representation predominantly focuses on discriminating instances generated from simple data augmentation schemes. However, the learned representation…