works on

From the 1 of 15 linked papers with an AI index.

most citedDynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition

1 citations · 1 across the 5 of their papers we have counts for

collaborators

15 papers

cs.CV2026

TreeSRNF: Square-Root Normal Fields for Generative Modelling of the Geometric and Structural Variability in Tree-like 3D Objects

Tahmina Khanam, Hamid Laga, Mohammed Bennamoun +4

The paper presents TreeSRNF, a mathematical framework that extends Square-Root Normal Fields to model both the geometry and branching structure of tree-like 3D objects, enabling an…

cs.CV2026

Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

Lian Xu, Mohammed Bennamoun, Farid Boussaid +3

Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily opera…

cs.CV2026

Hybrid Transformer-Mamba for Weakly Supervised Volumetric Medical Segmentation

Yiheng Lyu, Lian Xu, Coen Arrow +3

Weakly supervised segmentation enables model training from plane-level labels. Existing methods often rely on 2D encoders, neglecting the volumetric nature of medical data. We prop…

cs.CV2026

SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition

Yanan Liu, Anqi Zhu, Jingmin Zhu +6

Skeleton-based action recognition aims to understand human behaviors from body joint sequences and is especially challenging in the one-shot setting, where only a single labeled ex…

cs.CV2026

DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition

Jingmin Zhu, Anqi Zhu, James Bailey +5

Zero-shot skeleton-based action recognition (ZS-SAR) is fundamentally constrained by prevailing approaches that rely on aligning skeleton features with static, class-level semantic…

cs.CL2026

STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning

Zinuo Li, Yongxin Guo, Jun Liu +7

Human understanding of video dynamics relies on forming structured representations of entities, actions, and temporal relations before engaging in abstract reasoning. In contrast,…