3 citations · 5 across the 3 of their papers we have counts for
5 papers · 1 filter
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
Hector A. Valdez, Kyle Min, Subarna Tripathi
Pretraining egocentric vision-language models has become essential to improving downstream egocentric video-text tasks. These egocentric foundation models commonly use the transfor…
Contrastive Language Video Time Pre-training
Hengyue Liu, Kyle Min, Hector A. Valdez +1
We introduce LAVITI, a novel approach to learning language, video, and temporal representations in long-form videos via contrastive learning. Different from pre-training on video-t…
SViTT: Temporal Learning of Sparse Video-Text Transformers
Yi Li, Kyle Min, Subarna Tripathi +1
Do video-text transformers learn to model temporal relationships across frames? Despite their immense capacity and the abundance of multimodal training data, recent work has reveal…
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
Kyle Min, Sourya Roy, Subarna Tripathi +2
Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over…
Learning Spatial-Temporal Graphs for Active Speaker Detection
Sourya Roy, Kyle Min, Subarna Tripathi +2
We address the problem of active speaker detection through a new framework, called SPELL, that learns long-range multimodal graphs to encode the inter-modal relationship between au…