9 papers
Lower Bounds for Frank-Wolfe on Strongly Convex Sets
Jannis Halbey, Daniel Deza, Max Zimmer +3
We present a constructive lower bound of for Frank-Wolfe (FW) when both the objective and the constraint set are smooth and strongly convex, showing that…
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
Moritz Wagner, Christophe Roux, Max Zimmer +1
Post-training pruning can substantially reduce LLM inference costs, but it often degrades quality unless the remaining weights are adapted. Since global retraining is expensive at…
Curvature-Dependent Lower Bounds for Frank-Wolfe
Jannis Halbey, Christophe Roux, Sebastian Pokutta
The Frank-Wolfe algorithm achieves a convergence rate of for smooth convex optimization over compact convex domains, accelerating to when bo…
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
Max Zimmer, Nico Pelleriti, Christophe Roux +1
AI tools and agents are reshaping how researchers work, from proving theorems to training neural networks. Yet for many, it remains unclear how these tools fit into everyday resear…
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
Louis Schiekiera, Max Zimmer, Christophe Roux +2
We investigate the extent to which an LLM's hidden-state geometry can be recovered from its behavior in psycholinguistic experiments. Across eight instruction-tuned transformer mod…
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
Max Zimmer, Christophe Roux, Moritz Wagner +2
The resource requirements of neural networks can be significantly reduced through pruning - the removal of seemingly less important parameters. However, for LLMs, full retraining t…