papers

Publications (8)

cs.CV2025

Multi-entity Video Transformers for Fine-Grained Video Representation Learning

Matthew Walmer, Rose Kanjirathinkal, Kai Sheng Tai +3

The area of temporally fine-grained video representation learning focuses on generating frame-by-frame representations for temporally dense tasks, such as fine-grained action phase…

cs.CV2022

Dual-Key Multimodal Backdoors for Visual Question Answering

Matthew Walmer, Karan Sikka, Indranil Sur +2

The success of deep learning has enabled advances in multimodal tasks that require non-trivial fusion of multiple input domains. Although multimodal models have shown potential in…

cs.CV2026

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders

Matthew Walmer, Saksham Suri, Anirud Aggarwal +1

The space of task-agnostic feature upsampling has emerged as a promising area of research to efficiently create denser features from pre-trained visual backbones. These methods act…

cs.CV2023

TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models

Indranil Sur, Karan Sikka, Matthew Walmer +5

We present a Multimodal Backdoor Defense technique TIJO (Trigger Inversion using Joint Optimization). Recent work arXiv:2112.07668 has demonstrated successful backdoor attacks on m…

cs.CV2023

Teaching Matters: Investigating the Role of Supervision in Vision Transformers

Matthew Walmer, Saksham Suri, Kamal Gupta +1

Vision Transformers (ViTs) have gained significant popularity in recent years and have proliferated into many applications. However, their behavior under different learning paradig…

cs.CV2024

LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors

Saksham Suri, Matthew Walmer, Kamal Gupta +1

We present a simple self-supervised method to enhance the performance of ViT features for dense downstream tasks. Our Lightweight Feature Transform (LiFT) is a straightforward and…

cs.CV2025

Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition

Pulkit Kumar, Shuaiyi Huang, Matthew Walmer +2

Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking hav…

cs.CV2020

APRICOT: A Dataset of Physical Adversarial Attacks on Object Detection

Anneliese Braunegg, Amartya Chakraborty, Michael Krumdick +6

Physical adversarial attacks threaten to fool object detection systems, but reproducible research on the real-world effectiveness of physical patches and how to defend against them…