3 papers
cs.CL2024
SimCT: A Simple Consistency Test Protocol in LLMs Development Lifecycle
Fufangchen Zhao, Guoqiang Jin, Rui Zhao +2
In this work, we report our efforts to advance the standard operation procedure of developing Large Language Models (LLMs) or LLMs-based systems or services in industry. We introdu…
cs.CL2024
CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
Jiawei Gu, Zacc Yang, Chuanghao Ding +2
Large Language Models (LLMs) excel in diverse tasks but often underperform in specialized fields due to limited domain-specific or proprietary corpus. Continual pre-training (CPT)…
cs.CV2024
What Makes Good Few-shot Examples for Vision-Language Models?
Zhaojun Guo, Jinghui Lu, Xuejing Liu +3
Despite the notable advancements achieved by leveraging pre-trained vision-language (VL) models through few-shot tuning for downstream tasks, our detailed empirical study highlight…