collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

RegionReasoner: Region-Grounded Multi-Round Visual Reasoning

Wenfang Sun, Hao Chen, Yingjun Du +2

Large vision-language models have achieved remarkable progress in visual reasoning, yet most existing systems rely on single-step or text-only reasoning, limiting their ability to…

cs.CV2026

Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video

Zihui Gao, Ke Liu, Donny Y. Chen +4

Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet vid…

cs.CV2026

CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning

Wenxin Ma, Chenlong Wang, Ruisheng Yuan +6

Humans can look at a static scene and instantly predict what happens next -- will moving this object cause a collision? We call this ability Causal Spatial Reasoning. However, curr…

cs.CV2025

HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation

Cong Chen, Ziyuan Huang, Cheng Zou +6

In this work, we present HieraTok, a novel multi-scale Vision Transformer (ViT)-based tokenizer that overcomes the inherent limitation of modeling single-scale representations. Thi…

cs.CV2025

Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization

Canyu Zhao, Xiaoman Li, Tianjian Feng +3

We introduce Tinker, a versatile framework for high-fidelity 3D editing that operates in both one-shot and few-shot regimes without any per-scene finetuning. Unlike prior technique…

cs.CV2025

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Hao Zhong, Muzhi Zhu, Zongze Du +6

Long-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution fra…