Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
Makoto Shing, Kou Misaki, Han Bao +2
Causal language models have demonstrated remarkable capabilities, but their size poses significant challenges for deployment in resource-constrained environments. Knowledge distill…
cs.LG2023
Feature Normalization Prevents Collapse of Non-contrastive Learning Dynamics
Han Bao
Contrastive learning is a self-supervised representation learning framework, where two positive views generated through data augmentation are made similar by an attraction force in…