24 citations · 36 across the 10 of their papers we have counts for
17 papers · 1 filter
Thinking with Anchors: Grounded and Efficient Document Reasoning
Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13
Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…
MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation
Najmeh Sadoughi, Xinyu Li, Avijit Vajpayee +5
Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal a…
Large Scale Real-World Multi-Person Tracking
Bing Shuai, Alessandro Bergamo, Uta Buechler +3
This paper presents a new large scale multi-person tracking dataset -- \texttt{PersonPath22}, which is over an order of magnitude larger than currently available high quality multi…
An In-depth Study of Stochastic Backpropagation
Jun Fang, Mingze Xu, Hao Chen +3
In this paper, we provide an in-depth study of Stochastic Backpropagation (SBP) when training deep neural networks for standard image classification and object detection tasks. Dur…
Transfer of Representations to Video Label Propagation: Implementation Factors Matter
Daniel McKee, Zitong Zhan, Bing Shuai +3
This work studies feature representations for dense label propagation in video, with a focus on recently proposed methods that learn video correspondence using self-supervised sign…
Multi-Object Tracking with Hallucinated and Unlabeled Videos
Daniel McKee, Bing Shuai, Andrew Berneshawi +4
In this paper, we explore learning end-to-end deep neural trackers without tracking annotations. This is important as large-scale training data is essential for training deep neura…