activity
20212026
most citedScissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time

11 citations · 13 across the 13 of their papers we have counts for

collaborators

16 papers

cs.LG2026

AdaPaD: Adaptive Parallel Deflation for PEFT with Self-Correcting Rank Discovery

Barbara Su, Fangshuo Liao, Anastasios Kyrillidis

Fine-tuning large language models with LoRA requires choosing a rank r before training starts. Existing approaches either extract rank-1 components sequentially, freezing each comp…

cs.LG2026

SGD at the Edge of Stability: The Stochastic Sharpness Gap

Fangshuo Liao, Afroditi Kolomvaki, Anastasios Kyrillidis

When training neural networks with full-batch gradient descent (GD) and step size , the largest eigenvalue of the Hessian -- the sharpness -- rises to an…

cs.LG2026

Convergence Analysis of Two-Layer Neural Networks under Gaussian Input Masking

Afroditi Kolomvaki, Fangshuo Liao, Evan Dramko +2

We investigate the convergence guarantee of two-layer neural network training with Gaussian randomly masked inputs. This scenario corresponds to Gaussian dropout at the input level…

cs.DS2026

Exploiting Low-Rank Objective Structure in Discrete Quadratic Optimization

Ria Stevens, Fangshuo Liao, Barbara Su +3

We study the problem of maximizing a complex-valued quadratic form over the roots of unity. We show that when the objective matrix $\mathbf{Q}^\star \in \mathbb{C}^…

cs.LG2025

Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts

Fangshuo Liao, Anastasios Kyrillidis

Mixture-of-Experts (MoE) architectures have emerged as a cornerstone of modern AI systems. In particular, MoEs route inputs dynamically to specialized experts whose outputs are agg…

cs.LG2025

One Rank at a Time: Cascading Error Dynamics in Sequential Learning

Mahtab Alizadeh Vandchali, Fangshuo, Liao +1

Sequential learning -- where complex tasks are broken down into simpler, hierarchical components -- has emerged as a paradigm in AI. This paper views sequential learning through th…