17 citations · 61 across the 17 of their papers we have counts for
15 papers · 1 filter
Exploring Discrete Diffusion Models for Image Captioning
Zixin Zhu, Yixuan Wei, Jianfeng Wang +7
The image captioning task is typically realized by an auto-regressive method that decodes the text tokens one by one. We present a diffusion-based captioning model, dubbed the name…
Social Interpretable Tree for Pedestrian Trajectory Prediction
Liushuai Shi, Le Wang, Chengjiang Long +4
Understanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social I…
GeneAnnotator: A Semi-automatic Annotation Tool for Visual Scene Graph
Zhixuan Zhang, Chi Zhang, Zhenning Niu +2
In this manuscript, we introduce a semi-automatic scene graph annotation tool for images, the GeneAnnotator. This software allows human annotators to describe the existing relation…
Enriching Local and Global Contexts for Temporal Action Localization
Zixin Zhu, Wei Tang, Le Wang +2
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimin…
Adversarial Attack and Defense in Deep Ranking
Mo Zhou, Le Wang, Zhenxing Niu +3
Deep Neural Network classifiers are vulnerable to adversarial attack, where an imperceptible perturbation could result in misclassification. However, the vulnerability of DNN-based…
Video Imprint
Zhanning Gao, Le Wang, Nebojsa Jojic +3
A new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploit…