collaborators

11 papers

cs.CV2026

Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers

Ridma Jayasundara, Shaheer Mohamed, Tharindu Fernando +6

Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical system…

cs.CV2026

ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering

Paritosh Parmar, Eric Peh, Basura Fernando

Existing Causal-Why Video Question Answering (VideoQA) models often struggle with higher-order reasoning, relying on opaque, monolithic pipelines that entangle video understanding,…

cs.CV20262 cited

CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes

Paritosh Parmar, Eric Peh, Ruirui Chen +4

Causal video question answering (QA) has garnered increasing interest, yet existing datasets often lack depth in causal reasoning. To address this gap, we capitalize on the unique…

cs.CV2026

Learning to Visually Connect Actions and their Effects

Paritosh Parmar, Eric Peh, Basura Fernando

We introduce the novel concept of visually Connecting Actions and Their Effects (CATE) in video understanding. CATE can have applications in areas like task planning and learning f…

cs.CV2026

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning

Eric Peh, Debaditya Roy, Basura Fernando

Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining "black boxes", yet they remain prone to hallucination or rely on superficial shortcuts. I…

cs.CV2025

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

Xinyu Zhang, Yuxuan Dong, Lingling Zhang +5

Despite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input. When images contain large amounts of irrele…