most citedPartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-identification

4 citations · 11 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV20232 cited

DF-3DFace: One-to-Many Speech Synchronized 3D Face Animation with Diffusion

Se Jin Park, Joanna Hong, Minsu Kim +1

Speech-driven 3D facial animation has gained significant attention for its ability to create realistic and expressive facial animations in 3D space based on speech. Learning-based…

cs.CV2023

Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal Tokens

Minsu Kim, Jeongsoo Choi, Soumi Maiti +3

In this paper, we propose methods to build a powerful and efficient Image-to-Speech captioning (Im2Sp) model. To this end, we start with importing the rich knowledge related to ima…

cs.CV2023

Multi-Temporal Lip-Audio Memory for Visual Speech Recognition

Jeong Hun Yeo, Minsu Kim, Yong Man Ro

Visual Speech Recognition (VSR) is a task to predict a sentence or word from lip movements. Some works have been recently presented which use audio signals to supplement visual inf…

cs.CV20234 cited

PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-identification

Minsu Kim, Seungryong Kim, JungIn Park +2

Modern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data…

cs.CV2023

Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video

Minsu Kim, Chae Won Kim, Yong Man Ro

Forced alignment refers to a technology that time-aligns a given transcription with a corresponding speech. However, as the forced alignment technologies have developed using speec…

cs.CV20223 cited

Speaker-adaptive Lip Reading with User-dependent Padding

Minsu Kim, Hyunjun Kim, Yong Man Ro

Lip reading aims to predict speech based on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip ap…