6 papers
A Self-Attentive Meta-Optimizer with Group-Adaptive Learning Rates and Weight Decay
JiangBo Zhao, ZhaoXin Liu
Adaptive optimizers like AdamW apply uniform hyperparameters across all parameter groups, ignoring heterogeneous optimization dynamics across layers and modules. We address this li…
CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
Mengjun Yi, Hanwen Zhang, Hui Dou +2
Large pre-trained Vision-Language Models (VLMs), such as Contrastive Language-Image Pre-training (CLIP), have exhibited remarkable zero-shot performance across various image classi…
Enhancing Epidemic Forecasting: Evaluating the Role of Mobility Data and Graph Convolutional Networks
Suhan Guo, Zhenghao Xu, Furao Shen +1
Accurate prediction of contagious disease outbreaks is vital for informed decision-making. Our study addresses the gap between machine learning algorithms and their epidemiological…
SPAT: Sensitivity-based Multihead-attention Pruning on Time Series Forecasting Models
Suhan Guo, Jiahong Deng, Mengjun Yi +2
Attention-based architectures have achieved superior performance in multivariate time series forecasting but are computationally expensive. Techniques such as patching and adaptive…
Physics-inspired Energy Transition Neural Network for Sequence Learning
Zhou Wu, Junyi An, Baile Xu +2
Recently, the superior performance of Transformers has made them a more robust and scalable solution for sequence modeling than traditional recurrent neural networks (RNNs). Howeve…
Interactive Instance Annotation with Siamese Networks
Xiang Xu, Ruotong Li, Mengjun Yi +3
Annotating instance masks is time-consuming and labor-intensive. A promising solution is to predict contours using a deep learning model and then allow users to refine them. Howeve…