26 citations · 31 across the 5 of their papers we have counts for
7 papers
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
Yabing Wang, Le Wang, Qiang Zhou +4
Cross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during tr…
I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
Yiwei Ma, Jiayi Ji, Ke Ye +6
Significant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in t…
PolyBuilding: Polygon Transformer for End-to-End Building Extraction
Yuan Hu, Zhibin Wang, Zhou Huang +1
We present PolyBuilding, a fully end-to-end polygon Transformer for building extraction. PolyBuilding direct predicts vector representation of buildings from remote sensing images.…
Semantic Data Augmentation based Distance Metric Learning for Domain Generalization
Mengzhu Wang, Jianlong Yuan, Qi Qian +2
Domain generalization (DG) aims to learn a model on one or more different but related source domains that could be generalized into an unseen target domain. Existing DG methods try…
A Simple Baseline for Semi-supervised Semantic Segmentation with Strong Data Augmentation
Jianlong Yuan, Yifan Liu, Chunhua Shen +2
Recently, significant progress has been made on semantic segmentation. However, the success of supervised semantic segmentation typically relies on a large amount of labelled data,…
Instant-Teaching: An End-to-End Semi-Supervised Object Detection Framework
Qiang Zhou, Chaohui Yu, Zhibin Wang +2
Supervised learning based object detection frameworks demand plenty of laborious manual annotations, which may not be practical in real applications. Semi-supervised object detecti…