3 papers
cs.AR2026
PC2IM: An Efficient In-Memory Computing Accelerator for 3D Point Cloud
Dengfeng Wang, Shunqin Cai, Yanan Sun
3D point cloud neural networks have significantly enhanced the perceptual capabilities of resource-limited mobile intelligent systems. However, despite the transformative impact, t…
cs.AR2026
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
Xinyu Wang, Jieyu Li, Yanan Sun +1
Large Language Models (LLMs) incur substantial memory and computation costs. Prior works reduce FP-INT arithmetic overhead by converting linear-layer activations to block floating…
cs.AR2025
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
Zhican Wang, Hongxiang Fan, Haroon Waris +5
Large Language Models (LLMs) excel in natural language processing tasks but pose significant computational and memory challenges for edge deployment due to their intensive resource…