8 citations · 14 across the 4 of their papers we have counts for
4 papers · 1 filter
Collaborative Three-Stream Transformers for Video Captioning
Hao Wang, Libo Zhang, Heng Fan +1
As the most critical components in a sentence, subject, predicate and object require special attention in the video captioning task. To implement this idea, we design a novel frame…
AttMOT: Improving Multiple-Object Tracking by Introducing Auxiliary Pedestrian Attributes
Yunhao Li, Zhen Xiao, Lin Yang +4
Multi-object tracking (MOT) is a fundamental problem in computer vision with numerous applications, such as intelligent surveillance and automated driving. Despite the significant…
Deficiency-Aware Masked Transformer for Video Inpainting
Yongsheng Yu, Heng Fan, Libo Zhang
Recent video inpainting methods have made remarkable progress by utilizing explicit guidance, such as optical flow, to propagate cross-frame pixels. However, there are cases where…
PIDray: A Large-scale X-ray Benchmark for Real-World Prohibited Item Detection
Libo Zhang, Lutao Jiang, Ruyi Ji +1
Automatic security inspection relying on computer vision technology is a challenging task in real-world scenarios due to many factors, such as intra-class variance, class imbalance…