activity
20202025
most citedOccamy: A 432-Core Dual-Chiplet Dual-HBM2E 768-DP-GFLOP/s RISC-V System for 8-to-64-bit Dense and Sparse Computing in 12nm FinFET

13 citations · 26 across the 5 of their papers we have counts for

collaborators
Showing cs.ARShow all

7 papers · 1 filter

cs.AR20256 cited

Ramping Up Open-Source RISC-V Cores: Assessing the Energy Efficiency of Superscalar, Out-of-Order Execution

Zexin Fu, Riccardo Tedeschi, Gianmarco Ottavi +4

Open-source RISC-V cores are increasingly demanded in domains like automotive and space, where achieving high instructions per cycle (IPC) through superscalar and out-of-order (OoO…

cs.AR2025

CVA6S+: A Superscalar RISC-V Core with High-Throughput Memory Architecture

Riccardo Tedeschi, Gianmarco Ottavi, Côme Allart +11

Open-source RISC-V cores are increasingly adopted in high-end embedded domains such as automotive, where maximizing instructions per cycle (IPC) is becoming critical. Building on t…

cs.AR202513 cited

Occamy: A 432-Core Dual-Chiplet Dual-HBM2E 768-DP-GFLOP/s RISC-V System for 8-to-64-bit Dense and Sparse Computing in 12nm FinFET

Paul Scheffler, Thomas Benz, Viviane Potocnik +12

ML and HPC applications increasingly combine dense and sparse memory access computations to maximize storage efficiency. However, existing CPUs and GPUs struggle to flexibly handle…

cs.AR2024

Occamy: A 432-Core 28.1 DP-GFLOP/s/W 83% FPU Utilization Dual-Chiplet, Dual-HBM2E RISC-V-based Accelerator for Stencil and Sparse Linear Algebra Computations with 8-to-64-bit Floating-Point Support in 12nm FinFET

Gianna Paulin, Paul Scheffler, Thomas Benz +11

We present Occamy, a 432-core RISC-V dual-chiplet 2.5D system for efficient sparse linear algebra and stencil computations on FP64 and narrow (32-, 16-, 8-bit) SIMD FP data. Occamy…

cs.AR20222 cited

A Heterogeneous In-Memory Computing Cluster For Flexible End-to-End Inference of Real-World Deep Neural Networks

Angelo Garofalo, Gianmarco Ottavi, Francesco Conti +4

Deployment of modern TinyML tasks on small battery-constrained IoT devices requires high computational energy efficiency. Analog In-Memory Computing (IMC) using non-volatile memory…

cs.AR20215 cited

End-to-end 100-TOPS/W Inference With Analog In-Memory Computing: Are We There Yet?

Gianmarco Ottavi, Geethan Karunaratne, Francesco Conti +3

In-Memory Acceleration (IMA) promises major efficiency improvements in deep neural network (DNN) inference, but challenges remain in the integration of IMA within a digital system.…