23 papers
When Does Sparsity Mitigate the Curse of Depth in LLMs
Dilxat Muhtar, Xinyuan Song, Sebastian Pokutta +4
Recent work has demonstrated the curse of depth in large language models (LLMs), where later layers contribute less to learning and representation than earlier layers. Such under-u…
Lower Bounds for Frank-Wolfe on Strongly Convex Sets
Jannis Halbey, Daniel Deza, Max Zimmer +3
We present a constructive lower bound of for Frank-Wolfe (FW) when both the objective and the constraint set are smooth and strongly convex, showing that…
Neural Field Tokenizations with Hierarchy and Spatial Locality Priors
Alonso Urbano, David W. Romero, Max Zimmer +1
Neural fields parameterize data as functions from coordinates to values, providing a unified framework for representation learning across modalities. Existing approaches are domina…
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
Moritz Wagner, Christophe Roux, Max Zimmer +1
Post-training pruning can substantially reduce LLM inference costs, but it often degrades quality unless the remaining weights are adapted. Since global retraining is expensive at…
What Do Evolutionary Coding Agents Evolve?
Nico Pelleriti, Sree Harsha Nelaturu, Zhanke Zhou +4
Recent work pairs LLMs with evolutionary search to iteratively generate, modify, and select code using task-specific feedback. These systems have produced strong results in mathema…
RECON: Robust symmetry discovery via Explicit Canonical Orientation Normalization
Alonso Urbano, David W. Romero, Max Zimmer +1
Real world data often exhibits unknown, instance-specific symmetries that rarely exactly match a transformation group fixed a priori. Class-pose decompositions aim to create di…