3 papers
cs.LG2026
Residual Stream Analysis of Overfitting And Structural Disruptions
Quan Liu, Han Zhou, Wenquan Wu +2
Ensuring that large language models (LLMs) remain both helpful and harmless poses a significant challenge: fine-tuning on repetitive safety datasets, where unsafe prompts are paire…
cs.CL2025
AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment
Mengyu Bu, Shaolei Zhang, Zhongjun He +2
Multilingual large language models (LLMs) possess impressive multilingual understanding and generation capabilities. However, their performance and cross-lingual alignment often la…
cs.CL2025
Shall Your Data Strategy Work? Perform a Swift Study
Minlong Peng, Jingyi Yang, Zhongjun He +1
This work presents a swift method to assess the efficacy of particular types of instruction-tuning data, utilizing just a handful of probe examples and eliminating the need for mod…