activity
20162026
most citedI2DFormer: Learning Image to Document Attention for Zero-Shot Image Classification

21 citations · 28 across the 7 of their papers we have counts for

collaborators

12 papers

cs.CV2026

Elastic Token Compression for Pixel-Space Diffusion Transformers

Eduard Zamfir, Christian Reisswig, Zongwei Wu +2

Natural images concentrate their detail in a small fraction of the frame, yet diffusion models spend a full token on every patch, in every layer and at every timestep. The waste is…

cs.CV20222 cited

I2MVFormer: Large Language Model Generated Multi-View Document Supervision for Zero-Shot Image Classification

Muhammad Ferjad Naeem, Muhammad Gul Zain Ali Khan, Yongqin Xian +4

Recent works have shown that unstructured text (documents) from online sources can serve as useful auxiliary information for zero-shot image classification. However, these methods…

cs.CV202221 cited

I2DFormer: Learning Image to Document Attention for Zero-Shot Image Classification

Muhammad Ferjad Naeem, Yongqin Xian, Luc Van Gool +1

Despite the tremendous progress in zero-shot learning(ZSL), the majority of existing methods still rely on human-annotated attributes, which are difficult to annotate and scale. An…

cs.CV2022

Attribute Prototype Network for Any-Shot Learning

Wenjia Xu, Yongqin Xian, Jiuniu Wang +2

Any-shot image classification allows to recognize novel classes with only a few or even zero samples. For the task of zero-shot learning, visual attributes have been shown to play…

cs.CV20214 cited

Distilling Audio-Visual Knowledge by Compositional Contrastive Learning

Yanbei Chen, Yongqin Xian, A. Sophia Koepke +2

Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to t…

cs.CV2021

A Closer Look at Self-training for Zero-Label Semantic Segmentation

Giuseppe Pastore, Fabio Cermelli, Yongqin Xian +3

Being able to segment unseen classes not observed during training is an important technical challenge in deep learning, because of its potential to reduce the expensive annotation…