Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
In Tree Structure Should Sentence Be Generated
Yaguang Li, Xin Chen
Generative models reliant on sequential autoregression have been at the forefront of language generation for an extensive period, particularly following the introduction of widely…
cs.CL2024
Learning to Maximize Mutual Information for Chain-of-Thought Distillation
Xin Chen, Hanxian Huang, Yanjun Gao +3
Knowledge distillation, the technique of transferring knowledge from large, complex models to smaller ones, marks a pivotal step towards efficient AI deployment. Distilling Step-by…