1 citations · 4 across the 7 of their papers we have counts for
7 papers · 1 filter
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos
Zijia Lu, A S M Iftekhar, Gaurav Mittal +6
Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach ta…
Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment
Minh-Quan Le, Gaurav Mittal, Tianjian Meng +5
While diffusion models are powerful in generating high-quality, diverse synthetic data for object-centric tasks, existing methods struggle with scene-aware tasks such as Visual Que…
LOCL: Learning Object-Attribute Composition using Localization
Satish Kumar, ASM Iftekhar, Ekta Prashnani +1
This paper describes LOCL (Learning Object Attribute Composition using Localization) that generalizes composition zero shot learning to objects in cluttered and more realistic sett…
What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactions
A S M Iftekhar, Hao Chen, Kaustav Kundu +3
We propose a novel one-stage Transformer-based semantic and spatial refined transformer (SSRT) to solve the Human-Object Interaction detection task, which requires to localize huma…
StressNet: Detecting Stress in Thermal Videos
Satish Kumar, A S M Iftekhar, Michael Goebel +7
Precise measurement of physiological signals is critical for the effective monitoring of human vital signs. Recent developments in computer vision have demonstrated that signals su…
VSGNet: Spatial Attention Network for Detecting Human Object Interactions Using Graph Convolutions
Oytun Ulutan, A S M Iftekhar, B. S. Manjunath
Comprehensive visual understanding requires detection frameworks that can effectively learn and utilize object interactions while analyzing objects individually. This is the main o…