13 papers
Cross-Domain Acceleration of Open Modification Search: From Commodity Platforms to Emerging Memory and Storage Devices
Sumukh Pinge, Chang Eun Song, Po-Kai Hsu +10
Open modification search (OMS) in mass spectrometry (MS) is a data-intensive workload whose performance is dominantly limited by reference data movement rather than computation. Pr…
D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector Adaptation
Chang Eun Song, Sumukh Pinge, Tianqi Zhang +3
Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases. Howeve…
PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM Inference
Runyang Tian, Yanru Chen, Weihong Xu +1
Large language models are increasingly deployed on edge devices with tight power and area budgets. While mixed-precision GEMM reduces arithmetic complexity, quantized inference is…
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
Tsung-Han Lu, Zheyu Li, Minxuan Zhou +1
All-pairs shortest paths (APSP) is a fundamental algorithm used for routing, logistics, and network analysis, but the cubic time complexity and heavy data movement of the canonical…
FSL-HDnn: A 40 nm Few-shot On-Device Learning Accelerator with Integrated Feature Extraction and Hyperdimensional Computing
Weihong Xu, Chang Eun Song, Haichao Yang +5
This paper introduces FSL-HDnn, an energy-efficient accelerator that implements the end-to-end pipeline of feature extraction and on-device few-shot learning (FSL). The accelerator…
FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
You Hak Lee, Xiaofan Yu, Quanling Zhao +2
Unsupervised federated learning (UFL) has gained attention as a privacy-preserving, decentralized machine learning approach that eliminates the need for labor-intensive data labeli…