27 citations · 31 across the 3 of their papers we have counts for
3 papers
cs.CV2021
TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation
Tanzila Rahman, Mengyu Yang, Leonid Sigal
The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-moda…
cs.CV2021★ 4 cited
Mask-Guided Discovery of Semantic Manifolds in Generative Models
Mengyu Yang, David Rokeby, Xavier Snelgrove
Advances in the realm of Generative Adversarial Networks (GANs) have led to architectures capable of producing amazingly realistic images such as StyleGAN2, which, when trained on…
cs.HC2021★ 27 cited
Soloist: Generating Mixed-Initiative Tutorials from Existing Guitar Instructional Videos Through Audio Processing
Bryan Wang, Mengyu Yang, Tovi Grossman
Learning musical instruments using online instructional videos has become increasingly prevalent. However, pre-recorded videos lack the instantaneous feedback and personal tailorin…