activity
20242026
collaborators

13 papers

cs.CV2026

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

Kiet T. Nguyen, Hanbo Shim, Jinwoo Kim +1

Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spat…

cs.CV2026

When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models

Jiho Choi, Jaemin Kim, Sanghwan Kim +2

Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Visi…

cs.LG2026

Inverting Data Transformations via Diffusion Sampling

Jinwoo Kim, Sékou-Oumar Kaba, Jiyun Park +2

We study the problem of transformation inversion on general Lie groups: a datum is transformed by an unknown group element, and the goal is to recover an inverse transformation tha…

cs.LG2026

Flock: A Knowledge Graph Foundation Model via Learning on Random Walks

Jinwoo Kim, Xingyue Huang, Krzysztof Olejniczak +4

We study the problem of zero-shot link prediction on knowledge graphs (KGs), which requires models to generalize to novel entities and novel relations. Knowledge graph foundation m…

cs.LG2026

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

Yeonwoo Cha, Semin Kim, Jinhyeon Kwon +1

Any-to-any generation seeks to translate between arbitrary subsets of modalities, enabling flexible cross-modal synthesis. Despite recent success, existing flow-based approaches ar…

cs.CV2026

Training-Free Refinement of Flow Matching with Divergence-based Sampling

Yeonwoo Cha, Jaehoon Yoo, Semin Kim +3

Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to…