5 papers
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators
Fabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann
Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor AP…
BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal
Prabod Rathnayaka, Fabian Waschkowski, Lukas Wesemann
We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on this hardware to date. Existin…
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
Xinyu Tian, Shu Zou, Zhaoyuan Yang +5
Reasoning has emerged as a pivotal capability in Large Language Models (LLMs). Through Reinforcement Learning (RL), typically Group Relative Policy Optimization (GRPO), these model…
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
Shu Zou, Xinyu Tian, Lukas Wesemann +3
Prompting has emerged as a practical way to adapt frozen vision-language models (VLMs) for video anomaly detection (VAD). Yet, existing prompts are often overly abstract, overlooki…
Robust Spatiotemporal Forecasting Using Adaptive Deep-Unfolded Variational Mode Decomposition
Osama Ahmad, Lukas Wesemann, Fabian Waschkowski +1
Accurate spatiotemporal forecasting is critical for numerous complex systems but remains challenging due to complex volatility patterns and spectral entanglement in conventional gr…