11 citations · 13 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 2 cited
KeyPoint Relative Position Encoding for Face Recognition
Minchul Kim, Yiyang Su, Feng Liu +2
In this paper, we address the challenge of making ViT models more robust to unseen affine transformations. Such robustness becomes useful in various recognition tasks such as face…
cs.CV2023★ 11 cited
ChatGPT-Powered Hierarchical Comparisons for Image Classification
Zhiyuan Ren, Yiyang Su, Xiaoming Liu
The zero-shot open-vocabulary challenge in image classification is tackled by pretrained vision-language models like CLIP, which benefit from incorporating class-specific knowledge…
cs.CV2023
Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation
Yiyang Su, Ali Vosoughi, Shijian Deng +2
The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds la…