activity
20202024
most citedEmbedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning

24 citations · 40 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2024

Controllable Relation Disentanglement for Few-Shot Class-Incremental Learning

Yuan Zhou, Richang Hong, Yanrong Guo +3

In this paper, we propose to tackle Few-Shot Class-Incremental Learning (FSCIL) from a new perspective, i.e., relation disentanglement, which means enhancing FSCIL via disentanglin…

cs.CV2023★ 24 cited

Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning

Zijie Song, Zhenzhen Hu, Yuanen Zhou +3

Cross-lingual image captioning is a challenging task that requires addressing both cross-lingual and cross-modal obstacles in multimedia analysis. The crucial issue in this task is…

cs.CV2023

Advancing Incremental Few-shot Semantic Segmentation via Semantic-guided Relation Alignment and Adaptation

Yuan Zhou, Xin Chen, Yanrong Guo +3

Incremental few-shot semantic segmentation (IFSS) aims to incrementally extend a semantic segmentation model to novel classes according to only a few pixel-level annotated data, wh…

cs.CV2022★ 14 cited

Image Captioning via Compact Bidirectional Architecture

Zijie Song, Yuanen Zhou, Zhenzhen Hu +4

Most current image captioning models typically generate captions from left-to-right. This unidirectional property makes them can only leverage past context but not future context.…

cs.CV2021

Few-shot Learning with Global Relatedness Decoupled-Distillation

Yuan Zhou, Yanrong Guo, Shijie Hao +3

Despite the success that metric learning based approaches have achieved in few-shot learning, recent works reveal the ineffectiveness of their episodic training mode. In this paper…

cs.CV2021★ 2 cited

Semi-Autoregressive Transformer for Image Captioning

Yuanen Zhou, Yong Zhang, Zhenzhen Hu +1

Current state-of-the-art image captioning models adopt autoregressive decoders, \ie they generate each word by conditioning on previously generated words, which leads to heavy late…