9 citations · 13 across the 5 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2022★ 2 cited
Multimodal Tree Decoder for Table of Contents Extraction in Document Images
Pengfei Hu, Zhenrong Zhang, Jianshu Zhang +2
Table of contents (ToC) extraction aims to extract headings of different levels in documents to better understand the outline of the contents, which can be widely used for document…
cs.CV2022
Learning Audio-Visual embedding for Person Verification in the Wild
Peiwen Sun, Shanshan Zhang, Zishan Liu +4
It has already been observed that audio-visual embedding is more robust than uni-modality embedding for person verification. Here, we proposed a novel audio-visual strategy that co…
cs.CV2022★ 2 cited
Defensive Patches for Robust Recognition in the Physical World
Jiakai Wang, Zixin Yin, Pengfei Hu +5
To operate in real-world high-stakes environments, deep learning systems have to endure noises that have been continuously thwarting their robustness. Data-end defense, which impro…