14 papers
Scalable Attention for 5G NR Channel Estimation
Mahdi Abdollahpour, Marco Bertuletti, Yichao Zhang +2
Attention-based neural estimators achieve strong channel-estimation accuracy, but the computational cost of global attention over the time-frequency resource grid grows quadratical…
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
Yinrong Li, Zexin Fu, Yichao Zhang +5
Modern large language model workloads put increasing demands on parallel compute capability and on-chip memory capacity, while also stressing fine-grained data movement and synchro…
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
Marco Bertuletti, Yichao Zhang, Diyou Shen +3
The upcoming integration of AI in the physical layer (PHY) of 6G radio access networks (RAN) will enable a higher quality of service in challenging transmission scenarios. However,…
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
Yichao Zhang, Marco Bertuletti, Chi Zhang +5
Shared L1-memory clusters of streamlined instruction processors (processing elements - PEs) are commonly used as building blocks in modern, massively parallel computing architectur…
A Compute&Memory Efficient Model-Driven Neural 5G Receiver for Edge AI-assisted RAN
Mahdi Abdollahpour, Marco Bertuletti, Yichao Zhang +3
Artificial intelligence approaches for base-band processing for radio receivers have demonstrated significant performance gains. Most of the proposed methods are characterized by h…
A 410GFLOP/s, 64 RISC-V Cores, 204.8GBps Shared-Memory Cluster in 12nm FinFET with Systolic Execution Support for Efficient B5G/6G AI-Enhanced O-RAN
Yichao Zhang, Marco Bertuletti, Sergio Mazzola +2
We present HeartStream, a 64-RV-core shared-L1-memory cluster (410 GFLOP/s peak performance and 204.8 GBps L1 bandwidth) for energy-efficient AI-enhanced O-RAN. The cores and clust…