activity
20222024
most citedCAD -- Contextual Multi-modal Alignment for Dynamic AVQA

1 citations · 3 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2024

COSMU: Complete 3D human shape from monocular unconstrained images

Marco Pesavento, Marco Volino, Adrian Hilton

We present a novel framework to reconstruct complete 3D human shapes from a given target image by leveraging monocular unconstrained images. The objective of this work is to reprod…

cs.CV20241 cited

Benchmarking Monocular 3D Dog Pose Estimation Using In-The-Wild Motion Capture Data

Moira Shooter, Charles Malleson, Adrian Hilton

We introduce a new benchmark analysis focusing on 3D canine pose estimation from monocular in-the-wild images. A multi-modal dataset 3DDogs-Lab was captured indoors, featuring vari…

cs.CV2024

An Effective-Efficient Approach for Dense Multi-Label Action Detection

Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson +1

Unlike the sparse label action detection task, where a single action occurs in each timestamp of a video, in a dense multi-label scenario, actions can overlap. To address this chal…

cs.CV2024

ANIM: Accurate Neural Implicit Model for Human Reconstruction from a single RGB-D image

Marco Pesavento, Yuanlu Xu, Nikolaos Sarafianos +7

Recent progress in human shape learning, shows that neural implicit models are effective in generating 3D human surfaces from limited number of views, and even from a single RGB im…

cs.CV20231 cited

CAD -- Contextual Multi-modal Alignment for Dynamic AVQA

Asmar Nadeem, Adrian Hilton, Robert Dawes +2

In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA…

cs.CV2023

PAT: Position-Aware Transformer for Dense Multi-Label Action Detection

Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson +1

We present PAT, a transformer-based network that learns complex temporal co-occurrence action dependencies in a video by exploiting multi-scale temporal features. In existing metho…