6 papers
Post-Training Probability Manifold Correction via Structured SVD Pruning and Self-Referential Distillation
Aaron R. Flouro, Shawn P. Chadwick
Large language models are expensive to deploy. We introduce Sparse Knowledge Distillation (SparseKD), a post-training method that compresses transformer models by combining structu…
Adaptive Weighting in Knowledge Distillation: An Axiomatic Framework for Multi-Scale Teacher Ensemble Optimization
Aaron R. Flouro, Shawn P. Chadwick
Knowledge distillation with multiple teachers is increasingly used to improve robustness, efficiency, and safety, yet existing approaches rely largely on heuristic or implementatio…
Recursive Meta-Distillation: An Axiomatic Framework for Iterative Knowledge Refinement
Aaron R. Flouro, Shawn P. Chadwick
Recent work in probability-domain knowledge distillation has established axiomatic frameworks for temperature scaling, multi-teacher aggregation, and bias-variance trade-offs in si…
Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation
Aaron R. Flouro, Shawn P. Chadwick
Building on the probability-domain distillation framework of Sparse-KD, we develop an axiomatic, operator-theoretic framework for multi-teacher ensemble knowledge distillation. Rat…
Hallucinations Live in Variance
Aaron R. Flouro, Shawn P. Chadwick
Benchmarks measure whether a model is correct. They do not measure whether a model is reliable. This distinction is largely academic for single-shot inference, but becomes critical…
Sparse Knowledge Distillation: A Mathematical Framework for Probability-Domain Temperature Scaling and Multi-Stage Compression
Aaron R. Flouro, Shawn P. Chadwick
We develop a unified theoretical framework for sparse knowledge distillation based on probability-domain softening operators. While the equivalence $p^{1/T} \propto \mathrm{softmax…