6 papers
Thinking with Anchors: Grounded and Efficient Document Reasoning
Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13
Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…
WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling
Zelin Zhao, Min Shi, Bo Yuan +5
World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-Worl…
Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD Generation
Bo Yuan, Zelin Zhao, Petr Molodyk +2
Large language models have recently enabled text-to-CAD systems that synthesize parametric CAD programs (e.g., CadQuery) from natural-language prompts. In practice, however, geomet…
Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI
Lama Moukheiber, Caleb M. Yeung, Haotian Xue +3
Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without transparent reasoning or spatial e…
Laplacian Multi-scale Flow Matching for Generative Modeling
Zelin Zhao, Petr Molodyk, Haotian Xue +1
In this paper, we present Laplacian multiscale flow matching (LapFlow), a novel framework that enhances flow matching by leveraging multi-scale representations for image generative…
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
Zelin Zhao, Xinyu Gong, Bangya Liu +5
Achieving precise camera control in video generation remains challenging, as existing methods often rely on camera pose annotations that are difficult to scale to large and dynamic…