4 papers
MetaLogic: Robustness Evaluation of Text-to-Image Models via Logically Equivalent Prompts
Yifan Shen, Yangyang Shu, Hye-young Paik +1
Recent advances in text-to-image (T2I) models, especially diffusion-based architectures, have significantly improved the visual quality of generated images. However, these models c…
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
Wei Hua, Chenlin Zhou, Jibin Wu +2
The combination of Spiking Neural Networks (SNNs) with Vision Transformer architectures has garnered significant attention due to their potential for energy-efficient and high-perf…
PP-SSL : Priority-Perception Self-Supervised Learning for Fine-Grained Recognition
ShuaiHeng Li, Qing Cai, Fan Zhang +5
Self-supervised learning is emerging in fine-grained visual recognition with promising results. However, existing self-supervised learning methods are often susceptible to irreleva…
CIT: Rethinking Class-incremental Semantic Segmentation with a Class Independent Transformation
Jinchao Ge, Bowen Zhang, Akide Liu +4
Class-incremental semantic segmentation (CSS) requires that a model learn to segment new classes without forgetting how to segment previous ones: this is typically achieved by dist…