Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
Hardware Acceleration of Block-Diffusion LLM for Edge Devices
Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee +7
Single-stream (batch-one) edge inference cannot amortize weight traffic across requests. Full-attention diffusion LLMs recompute the entire sequence at every step; native block dif…
cs.AR2026
Thermal Tuning Overhead in Wafer-Scale Optical Interconnects for LLM MoE Training: A Cross-Layer Analysis and Ferroelectric-Based Mitigation
Seongwon Yoon, Pin-Jun Chen, Shimeng Yu
The rapid scaling of large language models (LLMs), particularly mixture-of-experts (MoE) architectures, has intensified interconnect demands because expert-parallel execution is co…