From the 1 of 19 linked papers with an AI index.
19 papers
CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking
Xiangqun Zhang, Likai Wang, Zekun Qian +2
The paper introduces CD-RMOT-Bench, a benchmark for evaluating how well referring multi-object tracking models trained on one visual domain perform on different, unseen domains, an…
Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms
Ruiqi Wu, Bingliang Jiao, Ruize Han +6
Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory…
Are All Tokens Necessary for Visual Place Recognition? An Empirical Study of Token Reduction for Efficient Inference
Tong Jin, Yunpeng Liu, Shuyu Hu +4
Recent visual place recognition (VPR) methods based on vision transformers, particularly foundation models, have achieved remarkable recognition performance. However, these models…
COVTrack++: Learning Open-Vocabulary Multi-Object Tracking from Continuous Videos via a Synergistic Paradigm
Zekun Qian, Wei Feng, Ruize Han +1
Multi-Object Tracking (MOT) has traditionally focused on a few specific categories, restricting its applicability to real-world scenarios involving diverse objects. Open-Vocabulary…
Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models
Yifeng Wu, Huimin Huang, Ruiluo Wu +5
Multimodal Large Language Models (MLLMs) often lose track of the right image regions during fine-grained spatial reasoning, because a textual query rarely carries any explicit geom…
IMWM: Intuition Models Complement World Models for Latent Planning
Baoqi Gao, Ruize Han, Miao Wang +1
Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this experimentally: even with a p…