7 papers
Imprint: Online Memory Compression for Long-Horizon Egocentric QA
Kousik Das, Debaditya Roy
Long-horizon egocentric question answering involves answering about events that have occurred hours or days in the past. This requires memory representations that remain both retri…
H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning
Eric Peh, Debaditya Roy, Basura Fernando
Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining "black boxes", yet they remain prone to hallucination or rely on superficial shortcuts. I…
Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation
Syed Wasiq, Syed Mohamad Tawseeq, Yashwant Pravinrao Bangde +1
Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reasoning remains largely unexplor…
Improving Temporal Action Segmentation via Constraint-Aware Decoding
Yeo Keat Ee, Debaditya Roy, Chen Li +2
Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challenges such as action variabil…
Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning
Yashwant Pravinrao Bangde, Debaditya Roy
Vision-Language Models (VLMs) exhibit strong performance in instruction following and open-ended vision-language reasoning, yet they frequently generate fluent outputs that are wea…
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
Dhruv Verma, Debaditya Roy, Basura Fernando
Situation recognition refers to the ability of an agent to identify and understand various situations or contexts based on available information and sensory inputs. It involves the…