Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
A Causal Perspective for Enhancing Jailbreak Attack and Defense
Licheng Pan, Yunsheng Lu, Jiexi Liu +5
Uncovering the mechanisms behind "jailbreaks" in large language models (LLMs) is crucial for enhancing their safety and reliability, yet these mechanisms remain poorly understood.…
cs.LG2025
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
Lei Liu, Hao Zhu, Yue Shen +4
Continual Pre-training (CPT) serves as a fundamental approach for adapting foundation models to domain-specific applications. Scaling laws for pre-training define a power-law relat…
cs.LG2025
Towards Real-world Debiasing: Rethinking Evaluation, Challenge, and Solution
Peng Kuang, Zhibo Wang, Zhixuan Chu +2
Spurious correlations in training data significantly hinder the generalization capability of machine learning models when faced with distribution shifts, leading to the proposition…