2 papers
cs.DC2026
Simulating Unified Tensor Resharding in heterogeneous AI systems
Sumit Kumar, Sayantan Dasgupta, Kushal Mitra +6
State-of-the-art AI training simulators assume homogeneous compute and network infrastructure. However, real-world training infrastructure is becoming increasingly heterogeneous si…
cs.CL2026
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
Sayantan Dasgupta, Trevor Cohn, Timothy Baldwin
The core learning signal used in language model distillation is the standard Kullback-Leibler (KL) divergence between the student and teacher distributions. Traditional KL divergen…