activity
20162024
most citedYour Negative May not Be True Negative: Boosting Image-Text Matching with False Negative Elimination

41 citations · 164 across the 35 of their papers we have counts for

collaborators
Showing cs.CVShow all

17 papers · 1 filter

cs.CV2024

Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge

Xiangyu Wu, Zhouyang Chi, Yang Yang +1

In this paper, we present our solution for the WSDM2023 Toloka Visual Question Answering Challenge. Inspired by the application of multimodal pre-trained models to various downstre…

cs.CV2024

The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA

Hailiang Zhang, Dian Chao, Zhihao Guan +1

In this paper, we introduce a grounded video question-answering solution. Our research reveals that the fixed official baseline method for video question answering involves two mai…

cs.CV2024

W-Net: A Facial Feature-Guided Face Super-Resolution Network

Hao Liu, Yang Yang, Yunxia Liu

Face Super-Resolution (FSR) aims to recover high-resolution (HR) face images from low-resolution (LR) ones. Despite the progress made by convolutional neural networks in FSR, the r…

cs.CV20241 cited

AMMUNet: Multi-Scale Attention Map Merging for Remote Sensing Image Segmentation

Yang Yang, Shunyi Zheng

The advancement of deep learning has driven notable progress in remote sensing semantic segmentation. Attention mechanisms, while enabling global modeling and utilizing contextual…

cs.CV20241 cited

Semi-Supervised Image Captioning Considering Wasserstein Graph Matching

Yang Yang

Image captioning can automatically generate captions for the given images, and the key challenge is to learn a mapping function from visual features to natural language features. E…

cs.CV2023

Solution for SMART-101 Challenge of ICCV Multi-modal Algorithmic Reasoning Task 2023

Xiangyu Wu, Yang Yang, Shengdong Xu +3

In this paper, we present our solution to a Multi-modal Algorithmic Reasoning Task: SMART-101 Challenge. Different from the traditional visual question-answering datasets, this cha…