2 papers
cs.CV2025
SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
Yipei Wang, Shiyu Hu, Shukun Jia +6
In this paper, we present the first systematic investigation and quantification of Similar Object Interference (SOI), a long-overlooked yet critical bottleneck in Single Object Tra…
cs.CV2025
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
Shiyu Hu, Xuchen Li, Xuzhao Li +4
Despite rapid progress in large vision-language models (LVLMs), existing video caption benchmarks remain limited in evaluating their alignment with human understanding. Most rely o…