1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2025
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
Arkaprava Sinha, Dominick Reilly, Francois Bremond +2
The introduction of vision-language models like CLIP has enabled the development of foundational video models capable of generalizing to unseen videos and human actions. However, t…
cs.CV2025
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
Arkaprava Sinha, Monish Soundar Raj, Pu Wang +3
Temporal Action Detection (TAD) in untrimmed videos poses significant challenges, particularly for Activities of Daily Living (ADL) requiring models to (1) process long-duration vi…
cs.CL2022★ 1 cited
MUG: Interactive Multimodal Grounding on User Interfaces
Tao Li, Gang Li, Jingjie Zheng +2
We present MUG, a novel interactive task for multimodal grounding where a user and an agent work collaboratively on an interface screen. Prior works modeled multimodal UI grounding…