collaborators

6 papers

cs.CV2026

Thinking with Anchors: Grounded and Efficient Document Reasoning

Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13

Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…

cs.CV2026

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

Zelin Zhao, Min Shi, Bo Yuan +5

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-Worl…

cs.LG2026

Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD Generation

Bo Yuan, Zelin Zhao, Petr Molodyk +2

Large language models have recently enabled text-to-CAD systems that synthesize parametric CAD programs (e.g., CadQuery) from natural-language prompts. In practice, however, geomet…

cs.CV2026

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI

Lama Moukheiber, Caleb M. Yeung, Haotian Xue +3

Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without transparent reasoning or spatial e…

cs.CV2026

Laplacian Multi-scale Flow Matching for Generative Modeling

Zelin Zhao, Petr Molodyk, Haotian Xue +1

In this paper, we present Laplacian multiscale flow matching (LapFlow), a novel framework that enhances flow matching by leveraging multi-scale representations for image generative…

cs.CV2025

CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization

Zelin Zhao, Xinyu Gong, Bangya Liu +5

Achieving precise camera control in video generation remains challenging, as existing methods often rely on camera pose annotations that are difficult to scale to large and dynamic…