Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models
Jiawei Xu, Minghui Liu, Aakriti Agrawal +2
Masked diffusion language models decode by iteratively unmasking tokens, where the unmasking order defines an "order of thought" that strongly influences generation quality yet is…
cs.LG2025
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters
Yiping Wang, Hanxian Huang, Yifang Chen +3
While Large language models (LLMs) have advanced natural language processing tasks, their growing computational and memory demands make deployment on resource-constrained devices l…
cs.LG2025
LeetDecoding: A PyTorch Library for Exponentially Decaying Causal Linear Attention with CUDA Implementations
Jiaping Wang, Simiao Zhang, Qiao-Chu He +1
The machine learning and data science community has made significant while dispersive progress in accelerating transformer-based large language models (LLMs), and one promising app…