3 papers
cs.AI2023
Talking Models: Distill Pre-trained Knowledge to Downstream Models via Interactive Communication
Zhe Zhao, Qingyun Liu, Huan Gui +3
Many recent breakthroughs in machine learning have been enabled by the pre-trained foundation models. By scaling up model parameters, training data, and computation resources, foun…
cs.LG2023
Create and Find Flatness: Building Flat Training Spaces in Advance for Continual Learning
Wenhang Shi, Yiren Chen, Zhe Zhao +3
Catastrophic forgetting remains a critical challenge in the field of continual learning, where neural networks struggle to retain prior knowledge while assimilating new information…
cs.CL2023
Recouple Event Field via Probabilistic Bias for Event Extraction
Xingyu Bai, Taiqiang Wu, Han Guo +7
Event Extraction (EE), aiming to identify and classify event triggers and arguments from event mentions, has benefited from pre-trained language models (PLMs). However, existing PL…