works on

From the 1 of 14 linked papers with an AI index.

most citedTreeGRNG: Binary Tree Gaussian Random Number Generator for Efficient Probabilistic AI Hardware

3 citations · 3 across the 2 of their papers we have counts for

collaborators

15 papers

cs.AR2026

ARES: Adaptive Reasoning-Effort Steering for PPA- and Cost-Aware RTL Optimization with LLM Agents

Stef Cuyckens, Mihaela Jivanescu, Jun Yin +2

The paper presents ARES, a framework that adaptively controls the reasoning effort of large language model agents when optimizing RTL designs for power, performance, and area, whil…

cs.AR2026

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

Chao Fang, Jun Yin, Man Shi +1

With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this ch…

cs.AR20263 cited

TreeGRNG: Binary Tree Gaussian Random Number Generator for Efficient Probabilistic AI Hardware

Jonas Crols, Guilherme Paim, Shirui Zhao +1

Bayesian Neural Networks (BNNs) offer opportunities for greatly enhancing the trustworthiness of conventional neural networks by monitoring the uncertainties in decision-making. A…

cs.AR2026

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

Yuzong Chen, Chao Fang, Xilai Dai +4

The substantial memory bandwidth and computational demands of large language models (LLMs) present critical challenges for efficient inference. To tackle this, the literature has e…

cs.AR2026

A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access

Xiaoling Yi, Ryan Antonio, Yunhao Deng +4

Achieving high compute utilization across a wide range of AI workloads is crucial for the efficiency of versatile DNN accelerators. This paper presents the Voltra chip and its util…

cs.AR2025

Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement

Yunhao Deng, Fanchen Kong, Xiaoling Yi +2

The growing disparity between computational power and on-chip communication bandwidth is a critical bottleneck in modern Systems-on-Chip (SoCs), especially for data-parallel worklo…