8 papers
Learning to Solve Generative ODEs Beyond the Linear Span
Sihyeon Kim, Seunghun Lee, Vikas Singh +1
Diffusion and flow generative models sample by integrating a learned ODE, but high quality still requires many sequential model evaluations. Solver learning reduces this cost by ad…
SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation
Hyunwoong Kim, Seongeun Lee, Hannah Yun +2
Group Relative Policy Optimization (GRPO) and its variants, originally developed for Large Language Models (LLMs), have recently been applied to Multimodal LLMs and produced strong…
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
Jihwan Park, Taehoon Song, Sanghyeok Lee +2
Vision-Language Models (VLMs) have been widely used in various visual recognition tasks due to their remarkable generalization capabilities. As these models grow in size and comple…
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
Jongha Kim, Minseong Bae, Sanghyeok Lee +2
Table images present unique challenges for effective and efficient understanding due to the need for question-specific focus and the presence of redundant background regions. Exist…
Representation Shift: Unifying Token Compression with FlashAttention
Joonmyung Choi, Sanghyeok Lee, Byungoh Ko +3
Transformers have demonstrated remarkable success across vision, language, and video. Yet, increasing task complexity has led to larger models and more tokens, raising the quadrati…
Efficient multi-view training for 3D Gaussian Splatting
Minhyuk Choi, Injae Kim, Hyunwoo J. Kim
3D Gaussian Splatting (3DGS) has emerged as a preferred choice alongside Neural Radiance Fields (NeRF) in inverse rendering due to its superior rendering speed. Currently, the comm…