3 citations · 5 across the 3 of their papers we have counts for
3 papers · 1 filter
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
Tanveer Hannan, Md Mohaiminul Islam, Jindong Gu +2
Large language models (LLMs) excel at retrieving information from lengthy text, but their vision-language counterparts (VLMs) face difficulties with hour-long videos, especially fo…
GRAtt-VIS: Gated Residual Attention for Auto Rectifying Video Instance Segmentation
Tanveer Hannan, Rajat Koner, Maximilian Bernhard +5
Recent trends in Video Instance Segmentation (VIS) have seen a growing reliance on online methods to model complex and lengthy video sequences. However, the degradation of represen…
InstanceFormer: An Online Video Instance Segmentation Framework
Rajat Koner, Tanveer Hannan, Suprosanna Shit +4
Recent transformer-based offline video instance segmentation (VIS) approaches achieve encouraging results and significantly outperform online approaches. However, their reliance on…