Showing cs.CVShow all
3 papers · 1 filter
cs.CV2021
Joint Learning of Visual-Audio Saliency Prediction and Sound Source Localization on Multi-face Videos
Minglang Qiao, Yufan Liu, Mai Xu +4
Visual and audio events simultaneously occur and both attract attention. However, most existing saliency prediction works ignore the influence of audio and only consider vision mod…
cs.CV2021
SDTP: Semantic-aware Decoupled Transformer Pyramid for Dense Image Prediction
Zekun Li, Yufan Liu, Bing Li +3
Although transformer has achieved great progress on computer vision tasks, the scale variation in dense image prediction is still the key challenge. Few effective multi-scale techn…
cs.CV2020
Towards Accurate Pixel-wise Object Tracking by Attention Retrieval
Zhipeng Zhang, Bing Li, Weiming Hu +1
The encoding of the target in object tracking moves from the coarse bounding-box to fine-grained segmentation map recently. Revisiting de facto real-time approaches that are capabl…