activity
20222026
most citedEvading Data Contamination Detection for Language Models is (too) Easy

3 citations · 4 across the 4 of their papers we have counts for

collaborators

7 papers

cs.SE2026

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

Thibaud Gloaguen, Niels Mündler, Mark Müller +2

A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although this practice is strongly encouraged by ag…

cs.LG2024

Gaussian Loss Smoothing Enables Certified Training with Tight Convex Relaxations

Stefan Balauca, Mark Niklas Müller, Yuhao Mao +3

Training neural networks with high certified accuracy against adversarial examples remains an open challenge despite significant efforts. While certification methods can effectivel…

cs.LG2024

SPEAR:Exact Gradient Inversion of Batches in Federated Learning

Dimitar I. Dimitrov, Maximilian Baader, Mark Niklas Müller +1

Federated learning is a framework for collaborative machine learning where clients only share gradient updates and not their private data with a server. However, it was recently sh…

cs.LG20243 cited

Evading Data Contamination Detection for Language Models is (too) Easy

Jasper Dekoninck, Mark Niklas Müller, Maximilian Baader +2

Large language models are widespread, with their performance on benchmarks frequently guiding user preferences for one model over another. However, the vast amount of data these mo…

cs.CV2024

Automated Classification of Model Errors on ImageNet

Momchil Peychev, Mark Niklas Müller, Marc Fischer +1

While the ImageNet dataset has been driving computer vision research over the past decade, significant label noise and ambiguity have made top-1 accuracy an insufficient measure of…

cs.CL2023

Prompt Sketching for Large Language Models

Luca Beurer-Kellner, Mark Niklas Müller, Marc Fischer +1

Many recent prompting strategies for large language models (LLMs) query the model multiple times sequentially -- first to produce intermediate results and then the final answer. Ho…