11 papers
Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction
Min-Jae Kim, Jun-Yeong Moon, Mujeen Sung +1
Backchannels, which signal listener states like empathy and understanding, are fundamental to natural human interaction. However, current approaches rely solely on audio and text.…
Detecting Unknown Objects via Energy-based Separation for Open World Object Detection
Jun-Woo Heo, Keonhee Park, Gyeong-Moon Park
In this work, we tackle the problem of Open World Object Detection (OWOD). This challenging scenario requires the detector to incrementally learn to classify known objects without…
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
Jongseo Lee, Wooil Lee, Gyeong-Moon Park +2
Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods based on…
ConceptSplit: Decoupled Multi-Concept Personalization of Diffusion Models via Token-wise Adaptation and Attention Disentanglement
Habin Lim, Yeongseob Won, Juwon Seo +1
In recent years, multi-concept personalization for text-to-image (T2I) diffusion models to represent several subjects in an image has gained much more attention. The main challenge…
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
Jongseo Lee, Kyungho Bae, Kyle Min +2
In this work, we tackle the problem of video classincremental learning (VCIL). Many existing VCIL methods mitigate catastrophic forgetting by rehearsal training with a few temporal…
When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series
Min-Yeong Park, Won-Jeong Lee, Seong Tae Kim +1
Recently, forecasting future abnormal events has emerged as an important scenario to tackle real-world necessities. However, the solution of predicting specific future time points…