5 citations · 8 across the 4 of their papers we have counts for
4 papers
Masked Autoencoder for Unsupervised Video Summarization
Minho Shim, Taeoh Kim, Jinhyung Kim +1
Summarizing a video requires a diverse understanding of the video, ranging from recognizing scenes to evaluating how much each frame is essential enough to be selected as a summary…
Decomposed Cross-modal Distillation for RGB-based Temporal Action Detection
Pilhyeon Lee, Taeoh Kim, Minho Shim +2
Temporal action detection aims to predict the time intervals and the classes of action instances in the video. Despite the promising performance, existing two-stream models exhibit…
You Only Train Once: Multi-Identity Free-Viewpoint Neural Human Rendering from Monocular Videos
Jaehyeok Kim, Dongyoon Wee, Dan Xu
We introduce You Only Train Once (YOTO), a dynamic human generation framework, which performs free-viewpoint rendering of different human identities with distinct motions, via only…
Exploring Temporally Dynamic Data Augmentation for Video Recognition
Taeoh Kim, Jinhyung Kim, Minho Shim +4
Data augmentation has recently emerged as an essential component of modern training recipes for visual recognition tasks. However, data augmentation for video recognition has been…