2 papers
cs.CV2025
PROFIT: A Specialized Optimizer for Deep Fine Tuning
Anirudh S Chakravarthy, Shuai Kyle Zheng, Xin Huang +4
The fine-tuning of pre-trained models has become ubiquitous in generative AI, computer vision, and robotics. Although much attention has been paid to improving the efficiency of fi…
cs.CL2024
The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis
Chen Yang, Junzhuo Li, Xinyao Niu +11
Uncovering early-stage metrics that reflect final model performance is one core principle for large-scale pretraining. The existing scaling law demonstrates the power-law correlati…