activity
20242026
most citedUnveiling the Statistical Foundations of Chain-of-Thought Prompting Methods

1 citations · 2 across the 7 of their papers we have counts for

collaborators

11 papers

cs.LG2026

On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking

Jianliang He, Leda Wang, Siyu Chen +1

We present a comprehensive analysis of how two-layer neural networks learn features to solve the modular addition task. Our work provides a full mechanistic interpretation of the l…

cs.AI2026

Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report

Zhuoran Yang, Ed Li, Jianliang He +18

We present Foundation-Sec-8B-Reasoning, the first open-source native reasoning model for cybersecurity. Built upon our previously released Foundation-Sec-8B base model (derived fro…

cs.LG2025

Unlocking Out-of-Distribution Generalization in Transformers via Recursive Latent Space Reasoning

Awni Altabaa, Siyu Chen, John Lafferty +1

Systematic, compositional generalization beyond the training distribution remains a core challenge in machine learning -- and a critical bottleneck for the emergent reasoning abili…

cs.LG2025

Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders

Siyu Chen, Heejune Sheen, Xuyuan Xiong +2

We study the challenge of achieving theoretically grounded feature recovery using Sparse Autoencoders (SAEs) for the interpretation of Large Language Models. Existing SAE training…

stat.ML2025

Quantile-Optimal Policy Learning under Unmeasured Confounding

Zhongren Chen, Siyu Chen, Zhengling Qi +2

We study quantile-optimal policy learning where the goal is to find a policy whose reward distribution has the largest -quantile for some . We focus on the offline…

cs.LG2025

In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention

Jianliang He, Xintian Pan, Siyu Chen +1

We study how multi-head softmax attention models are trained to perform in-context learning on linear data. Through extensive empirical experiments and rigorous theoretical analysi…