3 papers
cs.LG2025
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
Gabriel J. Perin, Runjin Chen, Xuxi Chen +3
Large Language Models (LLMs) have become indispensable in real-world applications. However, their widespread adoption raises significant safety concerns, particularly in responding…
cs.LG2025
Make Optimization Once and for All with Fine-grained Guidance
Mingjia Shi, Ruihan Lin, Xuxi Chen +8
Learning to Optimize (L2O) enhances optimization efficiency with integrated neural networks. L2O paradigms achieve great outcomes, e.g., refitting optimizer, generating unseen solu…
cs.CL2025
Extracting and Understanding the Superficial Knowledge in Alignment
Runjin Chen, Gabriel Jacob Perin, Xuxi Chen +5
Alignment of large language models (LLMs) with human values and preferences, often achieved through fine-tuning based on human feedback, is essential for ensuring safe and responsi…