10 papers
Show Me Examples: Inferring Visual Concepts from Image Sets
Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3
Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared co…
Normalizing Trajectory Models
Jiatao Gu, Tianrong Chen, Ying Shen +3
Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a few coarse transitions. Exis…
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
Ying Shen, Tianrong Chen, Yuan Gao +6
Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text-image seq…
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
Chen Huang, Xianhang Li, Vimal Thilak +2
Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature prediction. However with the inhe…
Normalizing Flows with Iterative Denoising
Tianrong Chen, Jiatao Gu, David Berthelot +2
Normalizing Flows (NFs) are a classical family of likelihood-based methods that have received revived attention. Recent efforts such as TARFlow have shown that NFs are capable of a…
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3
Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multip…