activity
20162024
most citedSpatiotemporal Residual Networks for Video Action Recognition

493 citations · 602 across the 19 of their papers we have counts for

collaborators
Showing cs.CVShow all

20 papers · 1 filter

cs.CV2024

Visual Concept Connectome (VCC): Open World Concept Discovery and their Interlayer Connections in Deep Models

Matthew Kowal, Richard P. Wildes, Konstantinos G. Derpanis

Understanding what deep network models capture in their learned representations is a fundamental challenge in computer vision. We present a new methodology to understanding such vi…

cs.CV2024

Selective, Interpretable, and Motion Consistent Privacy Attribute Obfuscation for Action Recognition

Filip Ilic, He Zhao, Thomas Pock +1

Concerns for the privacy of individuals captured in public imagery have led to privacy-preserving action recognition. Existing approaches often suffer from issues arising through o…

cs.CV2023★ 1 cited

Understanding Video Transformers for Segmentation: A Survey of Application and Interpretability

Rezaul Karim, Richard P. Wildes

Video segmentation encompasses a wide range of categories of problem formulation, e.g., object, scene, actor-action and multimodal video segmentation, for delineating task-specific…

cs.CV2023★ 1 cited

Multiscale Memory Comparator Transformer for Few-Shot Video Segmentation

Mennatullah Siam, Rezaul Karim, He Zhao +1

Few-shot video segmentation is the task of delineating a specific novel class in a query video using few labelled support images. Typical approaches compare support and query featu…

cs.CV2023★ 2 cited

StepFormer: Self-supervised Step Discovery and Localization in Instructional Videos

Nikita Dvornik, Isma Hadji, Ran Zhang +4

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, wi…

cs.CV2023★ 1 cited

MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer

Rezaul Karim, He Zhao, Richard P. Wildes +1

In this paper, we present an end-to-end trainable unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in video. The presented Multiscale Encode…