3 papers
cs.CV2025
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
Seokun Kang, Taehwan Kim
Video action recognition is a challenging but important task for understanding and discovering what the video does. However, acquiring annotations for a video is costly, and semi-s…
astro-ph.IM2024
Deeper, Sharper, Faster: Application of Efficient Transformer to Galaxy Image Restoration
Hyosun Park, Yongsik Jo, Seokun Kang +2
The Transformer architecture has revolutionized the field of deep learning over the past several years in diverse areas, including natural language processing, code generation, ima…
cs.MM2023
Sound of Story: Multi-modal Storytelling with Audio
Jaeyeon Bae, Seokhoon Jeong, Seokun Kang +4
Storytelling is multi-modal in the real world. When one tells a story, one may use all of the visualizations and sounds along with the story itself. However, prior studies on story…