2 citations · 2 across the 4 of their papers we have counts for
6 papers · 1 filter
Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment
Peng Jin, Hao Li, Zesen Cheng +5
Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local de…
TG-VQA: Ternary Game of Video Question Answering
Hao Li, Peng Jin, Zesen Cheng +5
Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions…
Multi-granularity Interaction Simulation for Unsupervised Interactive Segmentation
Kehan Li, Yian Zhao, Zhennan Wang +6
Interactive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and med…
Fuzzy Positive Learning for Semi-supervised Semantic Segmentation
Pengchong Qiao, Zhidan Wei, Yu Wang +6
Semi-supervised learning (SSL) essentially pursues class boundary exploration with less dependence on human annotations. Although typical attempts focus on ameliorating the inevita…
DPR-CAE: Capsule Autoencoder with Dynamic Part Representation for Image Parsing
Canqun Xiang, Zhennan Wang, Wenbin Zou +1
Parsing an image into a hierarchy of objects, parts, and relations is important and also challenging in many computer vision tasks. This paper proposes a simple and effective capsu…
PR Product: A Substitute for Inner Product in Neural Networks
Zhennan Wang, Wenbin Zou, Chen Xu
In this paper, we analyze the inner product of weight vector w and data vector x in neural networks from the perspective of vector orthogonal decomposition and prove that the direc…