2 papers
cs.LG2026
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
Shichang Zhang, Hongzhe Du, Jiaqi W. Ma +1
Modern AI systems are typically developed through multiple stages-pretraining, fine-tuning rounds, and subsequent adaptation or alignment, where each stage builds on the previous o…
cs.CL2025
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
Hongzhe Du, Weikai Li, Min Cai +5
Post-training is essential for the success of large language models (LLMs), transforming pre-trained base models into more useful and aligned post-trained models. While plenty of w…