5 papers
ELiTeFormer: An Efficient Transformer for FPGAs
Victor Agostinelli, Nicolas Bohm Agostini, Antonino Tumeo
Transformer blocks are prevalent in large language model (LLM) but present deployment challenges due to their challenging computational and memory demands. While prior work has typ…
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
Matthew Raffel, Victor Agostinelli, Lizhong Chen
This paper discusses the construction, fine-tuning, and deployment of BeaverTalk, a cascaded system for speech-to-text translation as part of the IWSLT 2025 simultaneous translatio…
ML For Hardware Design Interpretability: Challenges and Opportunities
Raymond Baartmans, Andrew Ensinger, Victor Agostinelli +1
The increasing size and complexity of machine learning (ML) models have driven the growing need for custom hardware accelerators capable of efficiently supporting ML workloads. How…
Swift: High-Performance Sparse Tensor Contraction for Scientific Applications
Andrew Ensinger, Gabriel Kulp, Victor Agostinelli +2
In scientific fields such as quantum computing, physics, chemistry, and machine learning, high dimensional data are typically represented using sparse tensors. Tensor contraction i…
Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation
Matthew Raffel, Victor Agostinelli, Lizhong Chen
Large language models (LLMs) have achieved state-of-the-art performance in various language processing tasks, motivating their adoption in simultaneous translation. Current fine-tu…