collaborators
Showing cs.ARShow all

6 papers · 1 filter

cs.AR2026

Why Do Prefetchers Fail? Let Agents Answer

Xiangfeng Sun, Ceyu Xu, Ningzhi Ai +3

Hardware prefetchers are crucial to processor performance, yet their design remains labor-intensive and expert-driven. Architects inspect execution and memory-access traces, identi…

cs.AR2026

Cache-Resident LLM Inference in GB-Scale Last-Level Caches

Wanning Zhang, Tongzhou Gu, Marco Canini +2

Large language model (LLM) inference is increasingly dominated by data movement across the memory hierarchy. Recent 3D-stacked cache technologies have enabled GB-scale last-level c…

cs.AR2026

ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses

Mengming Li, Chenlu Miao, Buqing Xu +7

Irregular memory accesses pose challenges for effective and efficient data prefetching. While temporal prefetchers have recently shown promise for irregular memory access patterns,…

cs.AR2026

VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization

Yipu Zhang, Jintao Cheng, Xingyu Liu +8

3D reconstruction and view synthesis are fundamental to AR/VR, robotics, and digital twins. The Visual Geometry Grounded Transformer (VGGT) enables strong feed-forward 3D reconstru…

cs.AR2025

A Scalable Architecture for Efficient Multi-bit Fully Homomorphic Encryption

Jiaao Ma, Ceyu Xu, Lisa Wu Wills

In the era of cloud computing, privacy-preserving computation offloading is crucial for safeguarding sensitive data. Fully Homomorphic Encryption (FHE) enables secure processing of…

cs.AR20235 cited

MasterRTL: A Pre-Synthesis PPA Estimation Framework for Any RTL Design

Wenji Fang, Yao Lu, Shang Liu +5

In modern VLSI design flow, the register-transfer level (RTL) stage is a critical point, where designers define precise design behavior with hardware description languages (HDLs) l…