4 papers
Architecture-driven Shift: towards a lightweight selector for capturing the trends of logit shift
Zhong Ye, Yu Hu, Ruilin Tang
Continual Learning (CL) is a practical paradigm to utilize power of deep pre-trained neural networks, but which pre-trained model has a better ability to balance ``Plasticity-Stabi…
BaseCal: Unsupervised Confidence Calibration via Base Model Signals
Hexiang Tan, Wanli Yang, Junwei Zhang +7
Reliable confidence is essential for trusting the outputs of LLMs, yet widely deployed post-trained LLMs (PoLLMs) typically compromise this trust with severe overconfidence. In con…
CombinationTS: A Modular Framework for Understanding Time-Series Forecasting Models
Xiaorui Wang, Fanda Fan, Chenxi Wang +9
Recent progress in time-series forecasting has led to rapidly increasing architectural complexity, yet many reported State-of-the-Art gains are statistically fragile or misattribut…
Fine-tuning Done Right in Model Editing
Wanli Yang, Rui Tang, Hongyu Zang +6
Fine-tuning, a foundational method for adapting large language models, has long been considered ineffective for model editing. Here, we challenge this belief, arguing that the repo…