10 papers
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models
Kiet T. Nguyen, Hanbo Shim, Jinwoo Kim +1
Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spat…
Self-conditioned Flow Map Language Models via Fixed-point Flows
Jaehoon Yoo, Wonjung Kim, Floor Eijkelboom +4
Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text by conditioning on its own denoising est…
Posterior Refinement: Fast Language Generation via Any-Order Flow Maps
Manan Agarwal, Sheel Shah, Chanhyuk Lee +6
Non-autoregressive generation offers a powerful paradigm for iterative refinement, allowing models to recursively critique, erase and regenerate arbitrary subsets of tokens. Howeve…
Inverting Data Transformations via Diffusion Sampling
Jinwoo Kim, Sékou-Oumar Kaba, Jiyun Park +2
We study the problem of transformation inversion on general Lie groups: a datum is transformed by an unknown group element, and the goal is to recover an inverse transformation tha…
Flow Map Language Models: One-step Language Modeling via Continuous Denoising
Chanhyuk Lee, Jaehoon Yoo, Manan Agarwal +6
Language models based on discrete diffusion have attracted widespread interest for their potential to provide faster generation than autoregressive models. Despite their promise, t…
MUX: Continuous Reasoning via Multiplexed Tokens
Ayhan Suleymanzade, Halil Alperen Gozeten, Michael Bronstein +2
Language models solve complex problems by articulating intermediate reasoning steps in natural language. While effective, this process is computationally bottlenecked: each reasoni…