12 papers
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
Bowen Wang, Chi Zhang, Diyou Shen +3
The paper introduces Ventaglio, a hardware extension and ISA support for vector processors that efficiently executes sparse tensor contractions in Transformer inference, achieving…
RTLScout: Joint Agentic Code and Synthesis Optimization for Efficient Digital Circuits
Felix Arnold, Ryan Amaudruz, Dimitrios Tsaras +2
We present RTLScout, an autonomous system that combines LLM-driven agentic design with circuit-level synthesis optimization and arithmetic architecture sweeps. An LLM agent iterati…
Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding
Konstantin Berestizshevsky, Renzo Andri, Lukas Cavigelli
We present Top-Theta (Top-) Attention, a training-free method for sparsifying transformer attention during inference. Our key insight is that static, per-head thresholds can be…
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
Chi Zhang, Luca Colagrande, Renzo Andri +1
Attention accounts for an increasingly dominant fraction of total computation during inference for mixture-of-experts (MoE) models, making efficient acceleration critical. Emerging…
GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read Mapping
Julien Eudine, Chu Li, Zhuo Cheng +11
Genome sequencing has become a central focus in computational biology. A genome study typically begins with sequencing, which produces millions to billions of short DNA fragments k…
GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units
Maxence Bouvier, Ryan Amaudruz, Felix Arnold +2
As AI workloads proliferate, optimizing arithmetic units is becoming increasingly important for reducing the footprint of digital systems. Conventional design flows, which often re…