2 papers
cs.CL2026
PADD: Path-Aligned Decompression Distillation for Non-Router Teacher to Guide MoE Student Learning
Xinyue Peng, Yi Qian, Jiaojiao Lin +2
As large language models (LLMs) continue to scale, it becomes increasingly challenging to grow model capacity under fixed computation budgets. We propose Path-Aligned Decompression…
cs.CL2025
LongCodeZip: Compress Long Context for Code Language Models
Yuling Shi, Yichun Qian, Hongyu Zhang +2
Code generation under long contexts is becoming increasingly critical as Large Language Models (LLMs) are required to reason over extensive information in the codebase. While recen…