2 papers
cs.AR2026
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
Shien Zhu, Gustavo Alonso
Online Transaction Processing (OLTP) is a classic application with a growing business. CPU-based OLTP has low lock serving efficiency. The main reason is that most locks are cold,…
cs.CL2025
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
Shien Zhu, Samuel Bohl, Robin Oester +1
Mixture-of-Experts (MoE) Large Language Models (LLMs) efficiently scale-up the model while keeping relatively low inference cost. As MoE models only activate part of the experts, r…