4 papers · 1 filter
SimCT: A Simple Consistency Test Protocol in LLMs Development Lifecycle
Fufangchen Zhao, Guoqiang Jin, Rui Zhao +2
In this work, we report our efforts to advance the standard operation procedure of developing Large Language Models (LLMs) or LLMs-based systems or services in industry. We introdu…
CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
Jiawei Gu, Zacc Yang, Chuanghao Ding +2
Large Language Models (LLMs) excel in diverse tasks but often underperform in specialized fields due to limited domain-specific or proprietary corpus. Continual pre-training (CPT)…
Balancing Speciality and Versatility: A Coarse to Fine Framework for Mitigating Catastrophic Forgetting in Large Language Models
Hengyuan Zhang, Yanru Wu, Dawei Li +4
Aligned Large Language Models (LLMs) showcase remarkable versatility, capable of handling diverse real-world tasks. Meanwhile, aligned LLMs are also expected to exhibit speciality,…
Consistency Matters: Explore LLMs Consistency From a Black-Box Perspective
Fufangchen Zhao, Guoqiang Jin, Jiaheng Huang +2
Nowadays both commercial and open-source academic LLM have become the mainstream models of NLP. However, there is still a lack of research on LLM consistency, meaning that througho…