13 papers
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models
Kiet T. Nguyen, Hanbo Shim, Jinwoo Kim +1
Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spat…
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
Jiho Choi, Jaemin Kim, Sanghwan Kim +2
Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Visi…
Inverting Data Transformations via Diffusion Sampling
Jinwoo Kim, Sékou-Oumar Kaba, Jiyun Park +2
We study the problem of transformation inversion on general Lie groups: a datum is transformed by an unknown group element, and the goal is to recover an inverse transformation tha…
Flock: A Knowledge Graph Foundation Model via Learning on Random Walks
Jinwoo Kim, Xingyue Huang, Krzysztof Olejniczak +4
We study the problem of zero-shot link prediction on knowledge graphs (KGs), which requires models to generalize to novel entities and novel relations. Knowledge graph foundation m…
FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows
Yeonwoo Cha, Semin Kim, Jinhyeon Kwon +1
Any-to-any generation seeks to translate between arbitrary subsets of modalities, enabling flexible cross-modal synthesis. Despite recent success, existing flow-based approaches ar…
Training-Free Refinement of Flow Matching with Divergence-based Sampling
Yeonwoo Cha, Jaehoon Yoo, Semin Kim +3
Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to…