collaborators

6 papers

cs.CL2026

Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks

Haoming Wen, Shi Chen, Qingyu Shi +4

Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised…

cs.LG2026

Propagation of Chaos in Contextual Flow Maps

Shi Chen, Zhengjiang Lin, Kaizhao Liu +1

We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps (CFMs): dynamical systems that evolve a…

cs.LG2026

Scaling Limits of Long-Context Transformers

Giuseppe Bruno, Shi Chen, Zhengjiang Lin +2

We study the long-context limit of softmax self-attention with a fixed query and a random context of i.i.d. keys on the sphere, viewing the inverse temperature as the sc…

cs.LG2026

Quantitative Clustering in Mean-Field Transformer Models

Shi Chen, Zhengjiang Lin, Yury Polyanskiy +1

The evolution of tokens through deep transformer models can be modeled as an interacting particle system that has been shown to exhibit an asymptotic clustering behavior akin to th…

cs.LG2025

Critical attention scaling in long-context transformers

Shi Chen, Zhengjiang Lin, Yury Polyanskiy +1

As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length increases…

cs.LG2025

Residual connections provably mitigate oversmoothing in graph neural networks

Ziang Chen, Zhengjiang Lin, Shi Chen +2

Graph neural networks (GNNs) have achieved remarkable empirical success in processing and representing graph-structured data across various domains. However, a significant challeng…