20 citations · 26 across the 6 of their papers we have counts for
10 papers · 1 filter
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
Shaunak Halbe, Bhagyashree Puranik, Jayakrishnan Unnikrishnan +3
Modern video retrieval systems are expected to handle diverse tasks ranging from corpus-level retrieval, fine-grained moment localization to flexible multimodal querying. Specializ…
VIDEOP2R: Video Understanding from Perception to Reasoning
Yifan Jiang, Yueying Wang, Rui Zhao +4
Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promising results on improving reasoning…
Modality Agnostic Efficient Long Range Encoder
Toufiq Parag, Ahmed Elgammal
The long-context capability of recent large transformer models can be surmised to rely on techniques such as attention/model parallelism, as well as hardware-level optimizations. W…
SISL:Self-Supervised Image Signature Learning for Splicing Detection and Localization
Susmit Agrawal, Prabhat Kumar, Siddharth Seth +3
Recent algorithms for image manipulation detection almost exclusively use deep network models. These approaches require either dense pixelwise groundtruth masks, camera ids, or ima…
Multilayer Dense Connections for Hierarchical Concept Classification
Toufiq Parag, Hongcheng Wang
Classification is a pivotal function for many computer vision tasks such as object classification, detection, scene segmentation. Multinomial logistic regression with a single fina…
VideoSSL: Semi-Supervised Learning for Video Classification
Longlong Jing, Toufiq Parag, Zhe Wu +2
We propose a semi-supervised learning approach for video classification, VideoSSL, using convolutional neural networks (CNN). Like other computer vision tasks, existing supervised…