11 papers
Show Me Examples: Inferring Visual Concepts from Image Sets
Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3
Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared co…
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
Ying Shen, Tianrong Chen, Yuan Gao +6
Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text-image seq…
Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
Ming Gui, Johannes Schusterbauer, Timy Phan +4
We introduce Representation Tokenizer (RepTok), a generative modeling framework that represents an image using a single continuous latent token obtained from self-supervised vision…
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3
Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multip…
SimpleFold: Folding Proteins is Simpler than You Think
Yuyang Wang, Jiarui Lu, Navdeep Jaitly +2
Protein folding models have achieved groundbreaking results typically via a combination of integrating domain knowledge into the architectural blocks and training pipelines. Noneth…
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
Jiatao Gu, Ying Shen, Tianrong Chen +6
Normalizing flows (NFs) are end-to-end likelihood-based generative models for continuous data, and have recently regained attention with encouraging progress on image generation. Y…