2 papers
cs.CV2025
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10
Jiangning Zhang, Junwei Zhu, Teng Hu +7
Native 4K (21603840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, ma…
cs.CL2025
LLM-Oriented Token-Adaptive Knowledge Distillation
Xurong Xie, Zhucun Xue, Jiafu Wu +5
Knowledge distillation (KD) is a key technique for compressing large-scale language models (LLMs), yet prevailing logit-based methods typically employ static strategies that are mi…