41 citations · 164 across the 35 of their papers we have counts for
17 papers · 1 filter
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
Xiangyu Wu, Zhouyang Chi, Yang Yang +1
In this paper, we present our solution for the WSDM2023 Toloka Visual Question Answering Challenge. Inspired by the application of multimodal pre-trained models to various downstre…
The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA
Hailiang Zhang, Dian Chao, Zhihao Guan +1
In this paper, we introduce a grounded video question-answering solution. Our research reveals that the fixed official baseline method for video question answering involves two mai…
W-Net: A Facial Feature-Guided Face Super-Resolution Network
Hao Liu, Yang Yang, Yunxia Liu
Face Super-Resolution (FSR) aims to recover high-resolution (HR) face images from low-resolution (LR) ones. Despite the progress made by convolutional neural networks in FSR, the r…
AMMUNet: Multi-Scale Attention Map Merging for Remote Sensing Image Segmentation
Yang Yang, Shunyi Zheng
The advancement of deep learning has driven notable progress in remote sensing semantic segmentation. Attention mechanisms, while enabling global modeling and utilizing contextual…
Semi-Supervised Image Captioning Considering Wasserstein Graph Matching
Yang Yang
Image captioning can automatically generate captions for the given images, and the key challenge is to learn a mapping function from visual features to natural language features. E…
Solution for SMART-101 Challenge of ICCV Multi-modal Algorithmic Reasoning Task 2023
Xiangyu Wu, Yang Yang, Shengdong Xu +3
In this paper, we present our solution to a Multi-modal Algorithmic Reasoning Task: SMART-101 Challenge. Different from the traditional visual question-answering datasets, this cha…