9 papers
DenseSteer: Steering Small Language Models towards Dense Math Reasoning
Yang Ouyang, Shuhang Lin, Jung-Eun Kim
Large language models (LLMs) demonstrate strong chain-of-thought (CoT) reasoning abilities, while smaller models (<= 3B parameters) significantly underperform on multi-step reasoni…
Position: Retire the "Positive Backdoor" Label -- Secret Alignment Requires Strict and Systematic Evaluation
Jianwei Li, Jung-Eun Kim
This position paper argues that the AI/ML community should stop overclaiming and retire the label "positive backdoor," and instead treat trigger-activated hidden behaviors as Secre…
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
Jianwei Li, Jung-Eun Kim
Backdoor attacks pose severe security threats to large language models (LLMs), where a model behaves normally under benign inputs but produces malicious outputs when a hidden trigg…
Superficial Safety Alignment Hypothesis
Jianwei Li, Jung-Eun Kim
As large language models (LLMs) are overwhelmingly more and more integrated into various applications, ensuring they generate safe responses is a pressing need. Previous studies on…
Impact of Layer Norm on Memorization and Generalization in Transformers
Rishi Singhal, Jung-Eun Kim
Layer Normalization (LayerNorm) is one of the fundamental components in transformers that stabilizes training and improves optimization. In recent times, Pre-LayerNorm transformers…
Explaining How Quantization Disparately Skews a Model
Abhimanyu Bellam, Jung-Eun Kim
Post Training Quantization (PTQ) is widely adopted due to its high compression capacity and speed with minimal impact on accuracy. However, we observed that disparate impacts are e…