6 papers
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
Jinho Park, Youbin Kim, Hogun Park +1
Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential ch…
JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
Hyunju Kang, Woohyun Lee, Jaewon Kim +1
Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, particularly through visually…
UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models
Hyunju Kang, Geonhee Han, Hogun Park
Node representation learning, such as Graph Neural Networks (GNNs), has emerged as a pivotal method in machine learning. The demand for reliable explanation generation surges, yet…
Balancing Graph Embedding Smoothness in Self-Supervised Learning via Information-Theoretic Decomposition
Heesoo Jung, Hogun Park
Self-supervised learning (SSL) in graphs has garnered significant attention, particularly in employing Graph Neural Networks (GNNs) with pretext tasks initially designed for other…
CIMAGE: Exploiting the Conditional Independence in Masked Graph Auto-encoders
Jongwon Park, Heesoo Jung, Hogun Park
Recent Self-Supervised Learning (SSL) methods encapsulating relational information via masking in Graph Neural Networks (GNNs) have shown promising performance. However, most exist…
AudioGenX: Explainability on Text-to-Audio Generative Models
Hyunju Kang, Geonhee Han, Yoonjae Jeong +1
Text-to-audio generation models (TAG) have achieved significant advances in generating audio conditioned on text descriptions. However, a critical challenge lies in the lack of tra…