14 citations · 14 across the 2 of their papers we have counts for
2 papers
cs.CV2025★ 14 cited
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
Young-Hu Park, Rae-Hong Park, Hyung-Min Park
This paper presents an efficient visual speech encoder for lip reading. While most recent lip reading studies have been based on the ResNet architecture and have achieved significa…
cs.LG2023
Unsupervised Speech Representation Pooling Using Vector Quantization
Jeongkyun Park, Kwanghee Choi, Hyunjun Heo +1
With the advent of general-purpose speech representations from large-scale self-supervised models, applying a single model to multiple downstream tasks is becoming a de-facto appro…