4 papers
WeGeFT: Weight-Generative Fine-Tuning for Multi-Faceted Efficient Adaptation of Large Models
Chinmay Savadikar, Xi Song, Tianfu Wu
Fine-tuning large pretrained Transformer models can focus on either introducing a small number of new learnable parameters (parameter efficiency) or editing representations of a sm…
PaCa-ViT: Learning Patch-to-Cluster Attention in Vision Transformers
Ryan Grainger, Thomas Paniagua, Xi Song +3
Vision Transformers (ViTs) are built on the assumption of treating image patches as ``visual tokens" and learn patch-to-patch attention. The patch embedding based tokenizer has a s…
AOGNets: Compositional Grammatical Architectures for Deep Learning
Xilai Li, Xi Song, Tianfu Wu
Neural architectures are the foundation for improving performance of deep neural networks (DNNs). This paper presents deep compositional grammatical architectures which harness the…
Towards Interpretable R-CNN by Unfolding Latent Structures
Tianfu Wu, Wei Sun, Xilai Li +2
This paper first proposes a method of formulating model interpretability in visual understanding tasks based on the idea of unfolding latent structures. It then presents a case stu…