2 papers
cs.CV2026
VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes
Yan Ma, Jiadi Su, Zhulin Hu +4
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Re…
cs.CV2026
Depth-Guided Video Object Counting in Crowded Scenes
Yuanjing Xu, Xinyan Liu, Weidong Chen +5
Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Exis…