40 citations · 122 across the 19 of their papers we have counts for
Showing 2021 · cs.CVShow all
2 papers · 2 filters
cs.CV2021★ 3 cited
MutualFormer: Multi-Modality Representation Learning via Cross-Diffusion Attention
Xixi Wang, Xiao Wang, Bo Jiang +2
Aggregating multi-modality data to obtain reliable data representation attracts more and more attention. Recent studies demonstrate that Transformer models usually work well for mu…
cs.CV2021★ 16 cited
Towards More Flexible and Accurate Object Tracking with Natural Language: Algorithms and Benchmark
Xiao Wang, Xiujun Shu, Zhipeng Zhang +4
Tracking by natural language specification is a new rising research topic that aims at locating the target object in the video sequence based on its language description. Compared…