4 papers
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
Aojie Jiang, Kang Zhu, Zhiheng Zhang +4
Tensor parallelism (TP) has become a key technique for latency-sensitive LLM inference, but it introduces frequent, tightly synchronized All-Reduce operations that lie directly on…
ElfCore: A 28nm Neural Processor Enabling Dynamic Structured Sparse Training and Online Self-Supervised Learning with Activity-Dependent Weight Update
Zhe Su, Giacomo Indiveri
In this paper, we present ElfCore, a 28nm digital spiking neural network processor tailored for event-driven sensory signal processing. ElfCore is the first to efficiently integrat…
An Efficient Multicast Addressing Encoding Scheme for Multi-Core Neuromorphic Processors
Zhe Su, Aron Bencsik, Giacomo Indiveri +1
Multi-core neuromorphic processors are becoming increasingly significant due to their energy-efficient local computing and scalable modular architecture, particularly for event-bas…
EchoSpike Predictive Plasticity: An Online Local Learning Rule for Spiking Neural Networks
Lars Graf, Zhe Su, Giacomo Indiveri
The drive to develop artificial neural networks that efficiently utilize resources has generated significant interest in bio-inspired Spiking Neural Networks (SNNs). These networks…