6 papers · 1 filter
Hardware Acceleration of Block-Diffusion LLM for Edge Devices
Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee +7
Single-stream (batch-one) edge inference cannot amortize weight traffic across requests. Full-attention diffusion LLMs recompute the entire sequence at every step; native block dif…
Thermal Tuning Overhead in Wafer-Scale Optical Interconnects for LLM MoE Training: A Cross-Layer Analysis and Ferroelectric-Based Mitigation
Seongwon Yoon, Pin-Jun Chen, Shimeng Yu
The rapid scaling of large language models (LLMs), particularly mixture-of-experts (MoE) architectures, has intensified interconnect demands because expert-parallel execution is co…
HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
Po-Kai Hsu, Weihong Xu, Qunyou Liu +2
Retrieval-Augmented Generation (RAG) relies on large-scale Approximate Nearest Neighbor Search (ANNS) to retrieve semantically relevant context for large language models. Among ANN…
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems
Wei-Hsing Huang, Jianwei Jia, Yuyao Kong +4
Recent developments have introduced Kolmogorov-Arnold Networks (KAN), an innovative architectural paradigm capable of replicating conventional deep neural network (DNN) capabilitie…
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
Wei-Hsing Huang, Janak Sharda, Cheng-Jhih Shih +6
Conventional large language models (LLMs) are equipped with dozens of GB to TB of model parameters, making inference highly energy-intensive and costly as all the weights need to b…
3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge Rendering
Wei-Hsing Huang, Cheng-Jhih Shih, Jian-Wei Su +10
Dynamic 3D Gaussian splatting (3DGS) extends static 3DGS to render dynamic scenes, enabling AR/VR applications with moving objects. However, implementing dynamic 3DGS on edge devic…