4 papers
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference
Junyi Luo, Xinting Jiang, Tai-Hao Wen +9
Microscaling (MX) is now the standard for low-bit large language model (LLM) inference. Its 4-bit form MXFP4 still loses substantial accuracy, because existing MX formats fix eithe…
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
Minghui Hou, Wei-Hsing Huang, Shaofeng Liang +5
Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonom…
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems
Wei-Hsing Huang, Jianwei Jia, Yuyao Kong +4
Recent developments have introduced Kolmogorov-Arnold Networks (KAN), an innovative architectural paradigm capable of replicating conventional deep neural network (DNN) capabilitie…
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
Wei-Hsing Huang, Jianwei Jia, Yuyao Kong +4
Recently, a novel model named Kolmogorov-Arnold Networks (KAN) has been proposed with the potential to achieve the functionality of traditional deep neural networks (DNNs) using or…