works on

From the 1 of 11 linked papers with an AI index.

activity
20242026
collaborators

11 papers

cs.CV2026

TreeSRNF: Square-Root Normal Fields for Generative Modelling of the Geometric and Structural Variability in Tree-like 3D Objects

Tahmina Khanam, Hamid Laga, Mohammed Bennamoun +4

The paper presents TreeSRNF, a mathematical framework that extends Square-Root Normal Fields to model both the geometry and branching structure of tree-like 3D objects, enabling an…

cs.CV2026

Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

Lian Xu, Mohammed Bennamoun, Farid Boussaid +3

Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily opera…

cs.CV2026

Hybrid Transformer-Mamba for Weakly Supervised Volumetric Medical Segmentation

Yiheng Lyu, Lian Xu, Coen Arrow +3

Weakly supervised segmentation enables model training from plane-level labels. Existing methods often rely on 2D encoders, neglecting the volumetric nature of medical data. We prop…

cs.CV2026

DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition

Jingmin Zhu, Anqi Zhu, James Bailey +5

Zero-shot skeleton-based action recognition (ZS-SAR) is fundamentally constrained by prevailing approaches that rely on aligning skeleton features with static, class-level semantic…

cs.CL2026

STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning

Zinuo Li, Yongxin Guo, Jun Liu +7

Human understanding of video dynamics relies on forming structured representations of entities, actions, and temporal relations before engaging in abstract reasoning. In contrast,…

cs.CV2026

Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection

A S M Sharifuzzaman Sagar, Mohammed Bennamoun, Farid Boussaid +4

In multimodal misinformation, deception usually arises not just from pixel-level manipulations in an image, but from the semantic and contextual claim jointly expressed by the imag…