6 papers
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
Haoming Wen, Shi Chen, Qingyu Shi +4
Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised…
Propagation of Chaos in Contextual Flow Maps
Shi Chen, Zhengjiang Lin, Kaizhao Liu +1
We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps (CFMs): dynamical systems that evolve a…
Scaling Limits of Long-Context Transformers
Giuseppe Bruno, Shi Chen, Zhengjiang Lin +2
We study the long-context limit of softmax self-attention with a fixed query and a random context of i.i.d. keys on the sphere, viewing the inverse temperature as the sc…
Quantitative Clustering in Mean-Field Transformer Models
Shi Chen, Zhengjiang Lin, Yury Polyanskiy +1
The evolution of tokens through deep transformer models can be modeled as an interacting particle system that has been shown to exhibit an asymptotic clustering behavior akin to th…
Critical attention scaling in long-context transformers
Shi Chen, Zhengjiang Lin, Yury Polyanskiy +1
As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length increases…
Residual connections provably mitigate oversmoothing in graph neural networks
Ziang Chen, Zhengjiang Lin, Shi Chen +2
Graph neural networks (GNNs) have achieved remarkable empirical success in processing and representing graph-structured data across various domains. However, a significant challeng…