47 citations · 47 across the 8 of their papers we have counts for
Showing 2025Show all
2 papers · 1 filter
cs.LG2025
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
Makoto Shing, Masanori Koyama, Takuya Akiba
End-to-end backpropagation requires storing activations throughout all layers, creating memory bottlenecks that limit model scalability. Existing block-wise training methods offer…
cs.LG2025
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
Makoto Shing, Kou Misaki, Han Bao +2
Causal language models have demonstrated remarkable capabilities, but their size poses significant challenges for deployment in resource-constrained environments. Knowledge distill…