2 papers
cs.CV2022
Panoramic Video Salient Object Detection with Ambisonic Audio Guidance
Xiang Li, Haoyuan Cao, Shijie Zhao +3
Video salient object detection (VSOD), as a fundamental computer vision problem, has been extensively discussed in the last decade. However, all existing works focus on addressing…
cs.CL2022
Relational Representation Learning in Visually-Rich Documents
Xin Li, Yan Zheng, Yiqing Hu +5
Relational understanding is critical for a number of visually-rich documents (VRDs) understanding tasks. Through multi-modal pre-training, recent studies provide comprehensive cont…