6 papers
Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse
Bole Ma, Jan Eitzinger, Harald Koestler +1
Multimodal agents repeatedly re-examine the same video frames, UI screenshots, and rendered artifacts as their context window slides and reasoning iterates, yet every look-back re-…
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
Martin Mayr, Sebastian Wind, Lukas Schröder +4
Artificial Intelligence (AI) workloads drive a rapid expansion of high-performance computing (HPC) infrastructures and increase their power and energy demands towards a critical le…
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
Bole Ma, Jan Eitzinger, Harald Köstler +1
Frontier LLMs increasingly decide what a query attends to with a sparse-attention indexer that picks a few KV-cache blocks per query: attention's unit is now a small, reusable chun…
Safety and accuracy follow different scaling laws in clinical large language models
Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa +9
Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation that higher accuracy implies…
SteuerLLM: Local specialized large language model for German tax law analysis
Sebastian Wind, Jeta Sopa, Laurin Schmid +8
Large language models (LLMs) demonstrate strong general reasoning and language understanding, yet their performance degrades in domains governed by strict formal rules, precise ter…
Multi-step retrieval and reasoning improves radiology question answering with large language models
Sebastian Wind, Jeta Sopa, Daniel Truhn +9
Clinical decision-making in radiology increasingly benefits from artificial intelligence (AI), particularly through large language models (LLMs). However, traditional retrieval-aug…