activity
20172023
most citedA Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation

99 citations · 186 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2022

Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis

Karren Yang, Dejan Markovic, Steven Krenn +2

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate…

cs.CV2022

LiP-Flow: Learning Inference-time Priors for Codec Avatars via Normalizing Flows in Latent Space

Emre Aksan, Shugao Ma, Akin Caliskan +5

Neural face avatars that are trained from multi-view data captured in camera domes can produce photo-realistic 3D reconstructions. However, at inference time, they must be driven b…

cs.CV2020

Audio- and Gaze-driven Facial Animation of Codec Avatars

Alexander Richard, Colin Lea, Shugao Ma +3

Codec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are al…

cs.CV201999 cited

A Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation

Hilde Kuehne, Alexander Richard, Juergen Gall

Action recognition has become a rapidly developing research field within the last decade. But with the increasing demand for large scale data, the need of hand annotated data for t…

cs.CV201911 cited

Mining YouTube - A dataset for learning fine-grained action concepts from webly supervised video data

Hilde Kuehne, Ahsan Iqbal, Alexander Richard +1

Action recognition is so far mainly focusing on the problem of classification of hand selected preclipped actions and reaching impressive results in this field. But with the perfor…

cs.CV2018

NeuralNetwork-Viterbi: A Framework for Weakly Supervised Video Learning

Alexander Richard, Hilde Kuehne, Ahsan Iqbal +1

Video learning is an important task in computer vision and has experienced increasing interest over the recent years. Since even a small amount of videos easily comprises several m…