activity
20242026
collaborators

6 papers

cs.CV2026

ChartSync: A Benchmark for Visuo-Logical Cascading Chart Editing

Jiakang Yu, Yixuan Chai, Tianci Wang +7

Generative image editing models struggle with structured statistical charts when data modifications require geometric synchronization. We formalize this task as Visuo-Logical Casca…

cs.CV2026

Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation

Guangkai Xu, Hua Geng, Huanyi Zheng +4

Feed-forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi-frame models usually produce better cross-frame consistency, yet…

cs.CV2025

Generative Video Matting

Yongtao Ge, Kangyang Xie, Guangkai Xu +6

Video matting has traditionally been limited by the lack of high-quality ground-truth data. Most existing video matting datasets provide only human-annotated imperfect alpha and fo…

eess.IV2025

POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction

Songyan Zhang, Yongtao Ge, Jinyuan Tian +4

3D reconstruction in dynamic scenes primarily relies on the combination of geometry estimation and matching modules where the latter task is pivotal for distinguishing dynamic regi…

cs.CV2024

What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

Guangkai Xu, Yongtao Ge, Mingyu Liu +5

Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent…

cs.CV2024

Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation

Muzhi Zhu, Yang Liu, Zekai Luo +5

The Diffusion Model has not only garnered noteworthy achievements in the realm of image generation but has also demonstrated its potential as an effective pretraining method utiliz…