12 citations · 12 across the 1 of their papers we have counts for
2 papers
cs.CV2022★ 12 cited
Exploring Discrete Diffusion Models for Image Captioning
Zixin Zhu, Yixuan Wei, Jianfeng Wang +7
The image captioning task is typically realized by an auto-regressive method that decodes the text tokens one by one. We present a diffusion-based captioning model, dubbed the name…
cs.CV2021
Enriching Local and Global Contexts for Temporal Action Localization
Zixin Zhu, Wei Tang, Le Wang +2
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimin…