4 citations · 6 across the 7 of their papers we have counts for
10 papers · 1 filter
TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding
Zhengwei Wang, Qi She, Aljosa Smolic
Most of existing video action recognition models ingest raw RGB frames. However, the raw video stream requires enormous storage and contains significant temporal redundancy. Video…
Foreground color prediction through inverse compositing
Sebastian Lutz, Aljosa Smolic
In natural image matting, the goal is to estimate the opacity of the foreground object in the image. This opacity controls the way the foreground and background is blended in trans…
DuctTake: Spatiotemporal Video Compositing
Jan Rueegg, Oliver Wang, Aljoscha Smolic +1
DuctTake is a system designed to enable practical compositing of multiple takes of a scene into a single video. Current industry solutions are based around object segmentation, a h…
CatNet: Class Incremental 3D ConvNets for Lifelong Egocentric Gesture Recognition
Zhengwei Wang, Qi She, Tejo Chalasani +1
Egocentric gestures are the most natural form of communication for humans to interact with wearable devices such as VR/AR helmets and glasses. A major issue in such scenarios for r…
Simultaneous Segmentation and Recognition: Towards more accurate Ego Gesture Recognition
Tejo Chalasani, Aljosa Smolic
Ego hand gestures can be used as an interface in AR and VR environments. While the context of an image is important for tasks like scene understanding, object recognition, image ca…
DublinCity: Annotated LiDAR Point Cloud and its Applications
S. M. Iman Zolanvari, Susana Ruano, Aakanksha Rana +4
Scene understanding of full-scale 3D models of an urban area remains a challenging task. While advanced computer vision techniques offer cost-effective approaches to analyse 3D urb…