collaborators

5 papers

cs.AR2026

PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM Inference

Runyang Tian, Yanru Chen, Weihong Xu +1

Large language models are increasingly deployed on edge devices with tight power and area budgets. While mixed-precision GEMM reduces arithmetic complexity, quantized inference is…

cs.AR2026

GEN-Graph: Heterogeneous PIM Accelerator for General Computational Patterns in Graph-based Dynamic Programming

Yanru Chen, Runyang Tian, Zheyu Li +3

While graph-based dynamic programming (DP) is a cornerstone of genomics and network analytics, its efficiency is hampered by fundamentally conflicting computational patterns. Matri…

cs.AR2025

CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference

Yanru Chen, Runyang Tian, Yue Pan +3

The proliferation of large language models (LLMs) is accelerating the integration of multimodal assistants into edge devices, where inference is executed under stringent latency an…

cs.AR2025

RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs

Yanru Chen, Zheyu Li, Keming Fan +5

All-pairs shortest paths (APSP) remains a major bottleneck for large-scale graph analytics, as data movement with cubic complexity overwhelms the bandwidth of conventional memory h…

cs.AR2025

HDDB: Efficient In-Storage SQL Database Search Using Hyperdimensional Computing on Ferroelectric NAND Flash

Quanling Zhao, Yanru Chen, Runyang Tian +6

Hyperdimensional Computing (HDC) encodes information and data into high-dimensional distributed vectors that can be manipulated using simple bitwise operations and similarity searc…