59 citations · 104 across the 10 of their papers we have counts for
12 papers · 1 filter
Training Strategies for Improved Lip-reading
Pingchuan Ma, Yujiang Wang, Stavros Petridis +2
Several training strategies and temporal models have been recently proposed for isolated word lip-reading in a series of independent works. However, the potential of combining the…
RISP: Rendering-Invariant State Predictor with Differentiable Simulation and Rendering for Cross-Domain Parameter Estimation
Pingchuan Ma, Tao Du, Joshua B. Tenenbaum +2
This work considers identifying parameters characterizing a physical system's dynamic motion directly from a video whose rendering configurations are inaccessible. Existing solutio…
Improving Deep Metric Learning by Divide and Conquer
Artsiom Sanakoyeu, Pingchuan Ma, Vadim Tschernezki +1
Deep metric learning (DML) is a cornerstone of many computer vision applications. It aims at learning a mapping from the input domain to an embedding space, where semantically simi…
End-to-end Audio-visual Speech Recognition with Conformers
Pingchuan Ma, Stavros Petridis, Maja Pantic
In this work, we present a hybrid CTC/Attention model based on a ResNet-18 and Convolution-augmented transformer (Conformer), that can be trained in an end-to-end manner. In partic…
A Content Transformation Block For Image Style Transfer
Dmytro Kotovenko, Artsiom Sanakoyeu, Pingchuan Ma +2
Style transfer has recently received a lot of attention, since it allows to study fundamental challenges in image understanding and synthesis. Recent work has significantly improve…
Lipreading using Temporal Convolutional Networks
Brais Martinez, Pingchuan Ma, Stavros Petridis +1
Lip-reading has attracted a lot of research attention lately thanks to advances in deep learning. The current state-of-the-art model for recognition of isolated words in-the-wild c…