7 papers · 1 filter
MoTE: Mixture of Task Experts for Multi-Task Video Understanding
Muhammad Asad Ali, Umar Khan, Nadia Robertini +1
Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transforme…
Amortized Inverse Kinematics via Graph Attention for Real-Time Human Avatar Animation
Muhammad Saif Ullah Khan, Chen-Yu Wang, Tim Prokosch +3
Inverse kinematics (IK) is a core operation in animation, robotics, and biomechanics: given Cartesian constraints, recover joint rotations under a known kinematic tree. In many rea…
GHOST: Fast Category-agnostic Hand-Object Interaction Reconstruction from RGB Videos using Gaussian Splatting
Ahmed Tawfik Aboukhadra, Marcel Rogge, Nadia Robertini +4
Understanding realistic hand-object interactions from monocular RGB videos is essential for AR/VR, robotics, and embodied AI. Existing methods rely on category-specific templates o…
Fast-HaMeR: Boosting Hand Mesh Reconstruction using Knowledge Distillation
Hunain Ahmed Jillani, Ahmed Tawfik Aboukhadra, Ahmed Elhayek +3
Fast and accurate 3D hand reconstruction is essential for real-time applications in VR/AR, human-computer interaction, robotics, and healthcare. Most state-of-the-art methods rely…
SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking
Muhammad Saif Ullah Khan, Didier Stricker
Modeling spinal motion is fundamental to understanding human biomechanics, yet remains underexplored in computer vision due to the spine's complex multi-joint kinematics and the la…
SurgeoNet: Realtime 3D Pose Estimation of Articulated Surgical Instruments from Stereo Images using a Synthetically-trained Network
Ahmed Tawfik Aboukhadra, Nadia Robertini, Jameel Malik +3
Surgery monitoring in Mixed Reality (MR) environments has recently received substantial focus due to its importance in image-based decisions, skill assessment, and robot-assisted s…