2 papers
cs.LG2025
Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
Jucheng Shen, Gaurav Sarkar, Yeonju Ro +4
We present CadLLM, a training-free method to accelerate the inference throughput of diffusion-based LLMs (dLLMs). We first investigate the dynamic nature of token unmasking confide…
cs.LG2025
SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
Gaurav Sarkar, Jay Gala, Syed Affan Daimi +1
Prevailing activation functions such as Swish and GELU tend toward domain-specific optima, Swish was discovered via neural architecture search on vision benchmarks, while GELU domi…