activity
20202026
most citedcosFormer: Rethinking Softmax in Attention

65 citations · 123 across the 18 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2023

All-pairs Consistency Learning for Weakly Supervised Semantic Segmentation

Weixuan Sun, Yanhao Zhang, Zhen Qin +5

In this work, we propose a new transformer-based regularization to better localize objects for Weakly supervised semantic segmentation (WSSS). In image-level WSSS, Class Activation…

cs.CV2023★ 16 cited

An Alternative to WSSS? An Empirical Study of the Segment Anything Model (SAM) on Weakly-Supervised Semantic Segmentation Problems

Weixuan Sun, Zheyuan Liu, Yanhao Zhang +2

The Segment Anything Model (SAM) has demonstrated exceptional performance and versatility, making it a promising tool for various related tasks. In this report, we explore the appl…

cs.CV2023★ 6 cited

Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder

Zheyuan Liu, Weixuan Sun, Damien Teney +1

Composed image retrieval aims to find an image that best matches a given multi-modal user query consisting of a reference image and text pair. Existing methods commonly pre-compute…

cs.CV2023★ 2 cited

Learning Audio-Visual Source Localization via False Negative Aware Contrastive Learning

Weixuan Sun, Jiayi Zhang, Jianyuan Wang +6

Self-supervised audio-visual source localization aims to locate sound-source objects in video frames without extra annotations. Recent methods often approach this goal with the hel…

cs.CV2023★ 2 cited

Bi-directional Training for Composed Image Retrieval via Text Prompt Learning

Zheyuan Liu, Weixuan Sun, Yicong Hong +2

Composed image retrieval searches for a target image based on a multi-modal user query comprised of a reference image and modification text describing the desired changes. Existing…

cs.CV2023★ 12 cited

Audio-Visual Segmentation with Semantics

Jinxing Zhou, Xuyang Shen, Jianyuan Wang +8

We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame…