3 papers
cs.AR2026
Hardware Acceleration of Block-Diffusion LLM for Edge Devices
Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee +7
Single-stream (batch-one) edge inference cannot amortize weight traffic across requests. Full-attention diffusion LLMs recompute the entire sequence at every step; native block dif…
cs.AR2025
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
Wei-Hsing Huang, Janak Sharda, Cheng-Jhih Shih +6
Conventional large language models (LLMs) are equipped with dozens of GB to TB of model parameters, making inference highly energy-intensive and costly as all the weights need to b…
cs.AR2025
3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge Rendering
Wei-Hsing Huang, Cheng-Jhih Shih, Jian-Wei Su +10
Dynamic 3D Gaussian splatting (3DGS) extends static 3DGS to render dynamic scenes, enabling AR/VR applications with moving objects. However, implementing dynamic 3DGS on edge devic…