3 papers
cs.CV2024
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
Skanda Koppula, Ignacio Rocco, Yi Yang +5
We introduce a new benchmark, TAPVid-3D, for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D). While point tracking in two dimensions (TAP) has many benchmarks m…
cs.CV2024
Neural Clustering based Visual Representation Learning
Guikun Chen, Xia Li, Yi Yang +1
We investigate a fundamental aspect of machine vision: the measurement of features, by revisiting clustering, one of the most classic approaches in machine learning and data analys…
cs.CV2023
Kefa: A Knowledge Enhanced and Fine-grained Aligned Speaker for Navigation Instruction Generation
Haitian Zeng, Xiaohan Wang, Wenguan Wang +1
We introduce a novel speaker model \textsc{Kefa} for navigation instruction generation. The existing speaker models in Vision-and-Language Navigation suffer from the large domain g…