6 papers
Count Anything at Any Granularity
Chang Liu, Haoning Wu, Weidi Xie
Open-world object counting remains brittle: despite rapid advances in vision-language models (VLMs), reliably counting the objects a user intends is far from solved. We argue that…
Track-On2: Enhancing Online Point Tracking with Memory
Görkay Aydemir, Weidi Xie, Fatma Güney
In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across video frames under significant appearance changes, mot…
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
Yibin Yan, Jilan Xu, Shangzhe Di +2
Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. However, current vision foundation…
Real-World Point Tracking with Verifier-Guided Pseudo-Labeling
Görkay Aydemir, Fatma Güney, Weidi Xie
Models for long-term point tracking are typically trained on large synthetic datasets. The performance of these models degrades in real-world videos due to different characteristic…
Track-On: Transformer-based Online Point Tracking with Memory
Görkay Aydemir, Xiongyi Cai, Weidi Xie +1
In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across multiple frames in a video, despite changes in appeara…
Can Visual Foundation Models Achieve Long-term Point Tracking?
Görkay Aydemir, Weidi Xie, Fatma Güney
Large-scale vision foundation models have demonstrated remarkable success across various tasks, underscoring their robust generalization capabilities. While their proficiency in tw…