collaborators

5 papers

cs.LG2026

Regularize or Localize: When Training-Time KV-Cache Geometry Pays Under Quantization

Libo Sun, Po-Wei Harn, Zewei Zhang +2

We study whether \sigreg -- LeJEPA's anti-collapse objective -- can reshape representations during standard autoregressive language-model pretraining, and when the resulting geomet…

cs.LG2026

Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks

Libo Sun, Po-Wei Harn, Zewei Zhang +2

Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We study \textbf{retrieval-warmed energy-based r…

cs.LG2026

Minimal-Intervention KV Retention via Set-Conditioned Diversity

Libo Sun, Po-wei Harn, Peixiong He +1

KV-cache compression at small budgets is a crowded design space spanning cache representation, head-wise routing, compression cadence, decoding behavior, and within-budget scoring.…

cs.CV2026

When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing

Libo Sun, Po-wei Harn, Peixiong He +1

Mixture-of-Experts (MoE) networks promise favorable accuracy-compute trade-offs, yet practical vision deployments are hindered by expert collapse and limited end-to-end efficiency…

cs.LG2026

MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression

Libo Sun, Peixiong He, Po-Wei Harn +1

KV cache memory is the dominant bottleneck for long-context LLM inference. Existing compression methods each act on a single axis of the four-dimensional KV tensor -- token evictio…