collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes

Yuling Xi, Haokai Zhang, Muzhi Zhu +10

Metric reasoning is a critical and challenging task for Vision Language Models (VLMs), playing a pivotal role in embodied AI tasks such as robotic manipulation and autonomous navig…

cs.CV2026

Exploring Spatial Intelligence from a Generative Perspective

Muzhi Zhu, Shunyao Jiang, Huanyi Zheng +9

Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern gener…

cs.CV2025

SE360: Semantic Edit in 360 Panoramas via Hierarchical Data Construction

Haoyi Zhong, Fang-Lue Zhang, Andrew Chalmers +1

While instruction-based image editing is emerging, extending it to 360 panoramas introduces additional challenges. Existing methods often produce implausible results in bot…

cs.CV2025

Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality

Zekai Luo, Zongze Du, Zhouhang Zhu +7

Video face swapping is crucial in film and entertainment production, where achieving high fidelity and temporal consistency over long and complex video sequences remains a signific…

cs.CV2025

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Hao Zhong, Muzhi Zhu, Zongze Du +6

Long-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution fra…

cs.CV2025

ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

Muzhi Zhu, Hao Zhong, Canyu Zhao +9

Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of effic…