activity
20222024
most citedUnderstanding Self-attention Mechanism via Dynamical System Perspective

3 citations · 6 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels

Haoxian Ruan, Zhihua Xu, Zhijing Yang +3

Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since col…

cs.CV20241 cited

Dynamic Correlation Learning and Regularization for Multi-Label Confidence Calibration

Tianshui Chen, Weihang Wang, Tao Pu +4

Modern visual recognition models often display overconfidence due to their reliance on complex deep neural networks and one-hot target supervision, resulting in unreliable confiden…

cs.IR2024

Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local Minima

Shanshan Zhong, Zhongzhan Huang, Daifeng Li +3

Multimodal recommender systems utilize various types of information to model user preferences and item features, helping users discover items aligned with their interests. The inte…

cs.CV2023

ADASR: An Adversarial Auto-Augmentation Framework for Hyperspectral and Multispectral Data Fusion

Jinghui Qin, Lihuang Fang, Ruitao Lu +2

Deep learning-based hyperspectral image (HSI) super-resolution, which aims to generate high spatial resolution HSI (HR-HSI) by fusing hyperspectral image (HSI) and multispectral im…

cs.CV20233 cited

Understanding Self-attention Mechanism via Dynamical System Perspective

Zhongzhan Huang, Mingfu Liang, Jinghui Qin +2

The self-attention mechanism (SAM) is widely used in various fields of artificial intelligence and has successfully boosted the performance of different models. However, current ex…

cs.CV2023

LSAS: Lightweight Sub-attention Strategy for Alleviating Attention Bias Problem

Shanshan Zhong, Wushao Wen, Jinghui Qin +2

In computer vision, the performance of deep neural networks (DNNs) is highly related to the feature extraction ability, i.e., the ability to recognize and focus on key pixel region…