collaborators

12 papers

cs.AR2026

At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference

Bowen Wang, Chi Zhang, Diyou Shen +3

The paper introduces Ventaglio, a hardware extension and ISA support for vector processors that efficiently executes sparse tensor contractions in Transformer inference, achieving…

cs.AR2026

RTLScout: Joint Agentic Code and Synthesis Optimization for Efficient Digital Circuits

Felix Arnold, Ryan Amaudruz, Dimitrios Tsaras +2

We present RTLScout, an autonomous system that combines LLM-driven agentic design with circuit-level synthesis optimization and arithmetic architecture sweeps. An LLM agent iterati…

cs.CL2026

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding

Konstantin Berestizshevsky, Renzo Andri, Lukas Cavigelli

We present Top-Theta (Top-) Attention, a training-free method for sparsifying transformer attention during inference. Our key insight is that static, per-head thresholds can be…

cs.AR2026

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators

Chi Zhang, Luca Colagrande, Renzo Andri +1

Attention accounts for an increasingly dominant fraction of total computation during inference for mixture-of-experts (MoE) models, making efficient acceleration critical. Emerging…

cs.AR2026

GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read Mapping

Julien Eudine, Chu Li, Zhuo Cheng +11

Genome sequencing has become a central focus in computational biology. A genome study typically begins with sequencing, which produces millions to billions of short DNA fragments k…

cs.LG2025

GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units

Maxence Bouvier, Ryan Amaudruz, Felix Arnold +2

As AI workloads proliferate, optimizing arithmetic units is becoming increasingly important for reducing the footprint of digital systems. Conventional design flows, which often re…