Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
Tianyu Wu, Yu Yao, Zhenting Qi +7
Speculative decoding accelerates LLM inference by having a small drafter propose tokens that a larger target model verifies in parallel. Recent diffusion-based parallel drafters su…
cs.LG2025
Frankenstein Optimizer: Harnessing the Potential by Revisiting Optimization Tricks
Chia-Wei Hsu, Nien-Ti Tsou, Yu-Cheng Chen +2
Gradient-based optimization drives the unprecedented performance of modern deep neural network models across diverse applications. Adaptive algorithms have accelerated neural netwo…