most citedLearning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

113 citations · 147 across the 7 of their papers we have counts for

collaborators

7 papers

cs.LG20235 cited

SEGA: Structural Entropy Guided Anchor View for Graph Contrastive Learning

Junran Wu, Xueyuan Chen, Bowen Shi +2

In contrastive learning, the choice of ``view'' controls the information that the representation captures and influences the performance of the model. However, leading graph contra…

physics.optics2023

Highly-confined and tunable plasmonics based on two-dimensional solid-state defect lattices

Ali Ghorashi, Nicholas Rivera, Bowen Shi +4

Plasmons, collective excitations of electrons in solids, are associated with strongly confined electromagnetic fields, with wavelengths far below the wavelength of photons in free…

cs.CL20234 cited

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

Mohamed Anwar, Bowen Shi, Vedanuj Goswami +3

We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languag…

cs.CV20235 cited

Pose-Oriented Transformer with Uncertainty-Guided Refinement for 2D-to-3D Human Pose Estimation

Han Li, Bowen Shi, Wenrui Dai +7

There has been a recent surge of interest in introducing transformers to 3D human pose estimation (HPE) due to their powerful capabilities in modeling long-term dependencies. Howev…

cs.AI20234 cited

Visual Story Generation Based on Emotion and Keywords

Yuetian Chen, Ruohua Li, Bowen Shi +2

Automated visual story generation aims to produce stories with corresponding illustrations that exhibit coherence, progression, and adherence to characters' emotional development.…

cs.CL202216 cited

u-HuBERT: Unified Mixed-Modal Speech Pretraining And Zero-Shot Transfer to Unlabeled Modality

Wei-Ning Hsu, Bowen Shi

While audio-visual speech models can yield superior performance and robustness compared to audio-only models, their development and adoption are hindered by the lack of labeled and…