activity
20242026
collaborators

6 papers

cs.CV2026

Count Anything at Any Granularity

Chang Liu, Haoning Wu, Weidi Xie

Open-world object counting remains brittle: despite rapid advances in vision-language models (VLMs), reliably counting the objects a user intends is far from solved. We argue that…

cs.CV2026

Track-On2: Enhancing Online Point Tracking with Memory

Görkay Aydemir, Weidi Xie, Fatma Güney

In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across video frames under significant appearance changes, mot…

cs.CV2026

OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams

Yibin Yan, Jilan Xu, Shangzhe Di +2

Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. However, current vision foundation…

cs.CV2026

Real-World Point Tracking with Verifier-Guided Pseudo-Labeling

Görkay Aydemir, Fatma Güney, Weidi Xie

Models for long-term point tracking are typically trained on large synthetic datasets. The performance of these models degrades in real-world videos due to different characteristic…

cs.CV2025

Track-On: Transformer-based Online Point Tracking with Memory

Görkay Aydemir, Xiongyi Cai, Weidi Xie +1

In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across multiple frames in a video, despite changes in appeara…

cs.CV2024

Can Visual Foundation Models Achieve Long-term Point Tracking?

Görkay Aydemir, Weidi Xie, Fatma Güney

Large-scale vision foundation models have demonstrated remarkable success across various tasks, underscoring their robust generalization capabilities. While their proficiency in tw…