5 papers
Vision Transformers are Circulant Attention Learners
Dongchen Han, Tianyu Li, Ziyi Wang +1
The self-attention mechanism has been a key factor in the advancement of vision Transformers. However, its quadratic complexity imposes a heavy computational burden in high-resolut…
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
Yifan Pu, Jixuan Ying, Qixiu Li +7
Vision Transformers (ViTs) have become a universal backbone for both image recognition and image generation. Yet their Multi-Head Self-Attention (MHSA) layer still performs a quadr…
Hierarchical Dual-Head Model for Suicide Risk Assessment via MentalRoBERTa
Chang Yang, Ziyi Wang, Wangfeng Tan +3
Social media platforms have become important sources for identifying suicide risk, but automated detection systems face multiple challenges including severe class imbalance, tempor…
Endo3R: Unified Online Reconstruction from Dynamic Monocular Endoscopic Video
Jiaxin Guo, Wenzhen Dong, Tianyu Huang +5
Reconstructing 3D scenes from monocular surgical videos can enhance surgeon's perception and therefore plays a vital role in various computer-assisted surgery tasks. However, achie…
Safe-VAR: Safe Visual Autoregressive Model for Text-to-Image Generative Watermarking
Ziyi Wang, Songbai Tan, Gang Xu +5
With the success of autoregressive learning in large language models, it has become a dominant approach for text-to-image generation, offering high efficiency and visual quality. H…