3 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 3 cited
SViTT: Temporal Learning of Sparse Video-Text Transformers
Yi Li, Kyle Min, Subarna Tripathi +1
Do video-text transformers learn to model temporal relationships across frames? Despite their immense capacity and the abundance of multimodal training data, recent work has reveal…
cs.CV2022★ 2 cited
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
Kyle Min, Sourya Roy, Subarna Tripathi +2
Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over…
cs.CV2021
Learning Spatial-Temporal Graphs for Active Speaker Detection
Sourya Roy, Kyle Min, Subarna Tripathi +2
We address the problem of active speaker detection through a new framework, called SPELL, that learns long-range multimodal graphs to encode the inter-modal relationship between au…