3 papers
cs.LG2026
BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference
Xiaoyou Wu, Cheng-Jhih Shih, Binfei Ji +2
Diffusion language models (dLLMs) generate text by iteratively denoising multiple token positions in parallel, offering an attractive alternative to strictly autoregressive decodin…
cs.AR2025
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
Wei-Hsing Huang, Janak Sharda, Cheng-Jhih Shih +6
Conventional large language models (LLMs) are equipped with dozens of GB to TB of model parameters, making inference highly energy-intensive and costly as all the weights need to b…
cs.AR2025
3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge Rendering
Wei-Hsing Huang, Cheng-Jhih Shih, Jian-Wei Su +10
Dynamic 3D Gaussian splatting (3DGS) extends static 3DGS to render dynamic scenes, enabling AR/VR applications with moving objects. However, implementing dynamic 3DGS on edge devic…