works on

From the 1 of 17 linked papers with an AI index.

most citedAddressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

2 citations · 2 across the 2 of their papers we have counts for

collaborators

17 papers

cs.LG2026

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis

A main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the model can use less compute for "ea…

cs.CV2026

Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models

Maulik Chevli, Johannes Brandt, Rickmer Braren +2

Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models could assist this process by providing g…

cs.LG20262 cited

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Jiazhen Pan, Bailiang Jian, Paul Hager +19

The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…

cs.LG2026

Efficient numeracy in language models through single-token number embeddings

Linus Kreitner, Paul Hager, Jonathan Mengedoht +3

To drive progress in science and engineering, large language models (LLMs) must be able to process large amounts of numerical data and solve long calculations efficiently. This is…

cs.LG2026

How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models

Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis

We measure how much one recurrence is worth to a looped (depth-recurrent) transformer, in equivalent unique parameters. From an iso-depth pretraining sweep across recurrence counts…

cs.CL2026

Weighting What Matters: Boosting Sample Efficiency in Medical Report Generation via Token Reweighting

Alexander Weers, Daniel Rueckert, Martin J. Menten

Training vision-language models (VLMs) for medical report generation is often hindered by the scarcity of high-quality annotated data. This work evaluates the use of a weighted los…