3 papers
cs.LG2026
Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective
Tokio Kajitsuka, Ukyo Honda, Sho Takase
Chain-of-thought (CoT) distillation transfers reasoning behaviors from a strong teacher to a smaller student, but prior work reports a capacity gap: distillation may fail when the…
cs.CL2025
Natural Fingerprints of Large Language Models
Teppei Suzuki, Ryokan Ri, Sho Takase
Recent studies have shown that the outputs from large language models (LLMs) can often reveal the identity of their source model. While this is a natural consequence of LLMs modeli…
cs.LG2025
Scaling Laws for Upcycling Mixture-of-Experts Language Models
Seng Pei Liew, Takuya Kato, Sho Takase
Pretraining large language models (LLMs) is resource-intensive, often requiring months of training time even with high-end GPU clusters. There are two approaches of mitigating such…