3 papers
cs.CV2024
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
Rezaul Karim, He Zhao, Richard P. Wildes +1
In this paper, we present an end-to-end trainable unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in video. The presented Multiscale Encode…
cs.CV2024
Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
Matthew Kowal, Mennatullah Siam, Md Amirul Islam +3
There is limited understanding of the information captured by deep spatiotemporal models in their intermediate representations. For example, while evidence suggests that action rec…
cs.CV2024
Visual Concept Connectome (VCC): Open World Concept Discovery and their Interlayer Connections in Deep Models
Matthew Kowal, Richard P. Wildes, Konstantinos G. Derpanis
Understanding what deep network models capture in their learned representations is a fundamental challenge in computer vision. We present a new methodology to understanding such vi…