activity
20242026
most citedAstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy

1 citations · 1 across the 9 of their papers we have counts for

collaborators

10 papers

cs.CV2026

Show Me Examples: Inferring Visual Concepts from Image Sets

Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3

Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared co…

cs.CV2026

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

Ying Shen, Tianrong Chen, Yuan Gao +6

Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text-image seq…

cs.LG2026

Text-Conditional JEPA for Learning Semantically Rich Visual Representations

Chen Huang, Xianhang Li, Vimal Thilak +2

Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature prediction. However with the inhe…

cs.CV2026

Normalizing Flows with Iterative Denoising

Tianrong Chen, Jiatao Gu, David Berthelot +2

Normalizing Flows (NFs) are a classical family of likelihood-based methods that have received revived attention. Recent efforts such as TARFlow have shown that NFs are capable of a…

cs.CV2026

Learning Long-term Motion Embeddings for Efficient Kinematics Generation

Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3

Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multip…

cs.LG2026

The Coupling Within: Flow Matching via Distilled Normalizing Flows

David Berthelot, Tianrong Chen, Jiatao Gu +6

Flow models have rapidly become the go-to method for training and deploying large-scale generators, owing their success to inference-time flexibility via adjustable integration ste…