8 papers
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
Shresth Grover, Priyank Pathak, Akash Kumar +1
Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual decision-making relatively underexpl…
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
Alejandro Aparcedo, Akash Kumar, Aaryan Garg +5
Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity t…
RobustGait: Robustness Analysis for Appearance Based Gait Recognition
Reeshoon Sayera, Akash Kumar, Sirshapan Mitra +2
Appearance-based gait recognition have achieved strong performance on controlled datasets, yet systematic evaluation of its robustness to real-world corruptions and silhouette vari…
OmViD: Omni-supervised active learning for video action detection
Aayush Rana, Akash Kumar, Vibhav Vineet +1
Video action detection requires dense spatio-temporal annotations, which are both challenging and expensive to obtain. However, real-world videos often vary in difficulty and may n…
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
Akash Kumar, Ashlesha Kumar, Vibhav Vineet +1
Self-supervised learning has emerged as a powerful paradigm for label-free model pretraining, particularly in the video domain, where manual annotation is costly and time-intensive…
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
Aaryan Garg, Akash Kumar, Yogesh S Rawat
In this work we study Weakly Supervised Spatio-Temporal Video Grounding (WSTVG), a challenging task of localizing subjects spatio-temporally in videos using only textual queries an…