3 papers
cs.LG2025
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
Amir Ziashahabi, Yavuz Faruk Bakman, Duygu Nur Yaldiz +3
Speculative Decoding (SD) ensures that the output matches the target model's distribution exactly. However, we argue that this distribution matching requirement is too stringent an…
cs.NE2025
Regularizing Differentiable Architecture Search with Smooth Activation
Yanlin Zhou, Mostafa El-Khamy, Kee-Bong Song
Differentiable Architecture Search (DARTS) is an efficient Neural Architecture Search (NAS) method but suffers from robustness, generalization, and discrepancy issues. Many efforts…
cs.CV2025
Hardware-Friendly Static Quantization Method for Video Diffusion Transformers
Sanghyun Yi, Qingfeng Liu, Mostafa El-Khamy
Diffusion Transformers for video generation have gained significant research interest since the impressive performance of SORA. Efficient deployment of such generative-AI models on…