activity
20162026
most citedWhat to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactions

1 citations · 2 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance

Kaustav Kundu, Ritvik Shrivastava, Maxim Arap +13

We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \…

cs.CV2022★ 1 cited

What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactions

A S M Iftekhar, Hao Chen, Kaustav Kundu +3

We propose a novel one-stage Transformer-based semantic and spatial refined transformer (SSRT) to solve the Human-Object Interaction detection task, which requires to localize huma…

cs.CV2022★ 1 cited

Hierarchical Self-supervised Representation Learning for Movie Understanding

Fanyi Xiao, Kaustav Kundu, Joseph Tighe +1

Most self-supervised video representation learning approaches focus on action recognition. In contrast, in this paper we focus on self-supervised video learning for movie understan…

cs.CV2021

TubeR: Tubelet Transformer for Video Action Detection

Jiaojiao Zhao, Yanyi Zhang, Xinyu Li +10

We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an off-line actor detector or hand-designed ac…

cs.CV2020

Positive-Congruent Training: Towards Regression-Free Model Updates

Sijie Yan, Yuanjun Xiong, Kaustav Kundu +5

Reducing inconsistencies in the behavior of different versions of an AI system can be as important in practice as reducing its overall error. In image classification, sample-wise i…

cs.CV2018

SurfConv: Bridging 3D and 2D Convolution for RGBD Images

Hang Chu, Wei-Chiu Ma, Kaustav Kundu +2

We tackle the problem of using 3D information in convolutional neural networks for down-stream recognition tasks. Using depth as an additional channel alongside the RGB input has t…