3 papers
cs.CV2026
SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange
Jaemo Jeong, Junho Yoon, Hyunju Kim +1
Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training-free methods query new even…
cs.CV2025
DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
Junho Yoon, Jaemo Jung, Hyunju Kim +1
Aligning egocentric video with wearable sensors have shown promise for human action recognition, but face practical limitations in user discomfort, privacy concerns, and scalabilit…
cs.LG2025
GPU Memory Prediction for Multimodal Model Training
Jinwoo Jeong, Minchul Kang, Younghun Go +5
As deep learning models in agentic AI systems grow in scale and complexity, GPU memory requirements increase and often exceed the available GPU memory capacity, so that out-of-memo…