8 papers · 1 filter
Imprint: Online Memory Compression for Long-Horizon Egocentric QA
Kousik Das, Debaditya Roy
Long-horizon egocentric question answering involves answering about events that have occurred hours or days in the past. This requires memory representations that remain both retri…
H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning
Eric Peh, Debaditya Roy, Basura Fernando
Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining "black boxes", yet they remain prone to hallucination or rely on superficial shortcuts. I…
Improving Temporal Action Segmentation via Constraint-Aware Decoding
Yeo Keat Ee, Debaditya Roy, Chen Li +2
Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challenges such as action variabil…
Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning
Yashwant Pravinrao Bangde, Debaditya Roy
Vision-Language Models (VLMs) exhibit strong performance in instruction following and open-ended vision-language reasoning, yet they frequently generate fluent outputs that are wea…
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
Dhruv Verma, Debaditya Roy, Basura Fernando
Situation recognition refers to the ability of an agent to identify and understand various situations or contexts based on available information and sensory inputs. It involves the…
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
Ramanathan Rajendiran, Debaditya Roy, Basura Fernando
Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich…