activity
20242026
most citedWeight-Entanglement Meets Gradient-Based Neural Architecture Search

3 citations · 3 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG2026

SVD-Surgeon: Optimal Singular-Value Surgery for Large Language Model Compression

Mahmoud Safari, Frank Hutter

Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their deployment is constrained by substantial memory and compute requirements. Low-ra…

cs.LG20253 cited

Weight-Entanglement Meets Gradient-Based Neural Architecture Search

Rhea Sanjay Sukthanker, Arjun Krishnakumar, Mahmoud Safari +1

Weight sharing is a fundamental concept in neural architecture search (NAS), enabling gradient-based methods to explore cell-based architectural spaces significantly faster than tr…

cs.LG2025

Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning Dynamics

Indrashis Das, Mahmoud Safari, Steven Adriaensen +1

Activation functions are fundamental elements of deep learning architectures as they significantly influence training dynamics. ReLU, while widely used, is prone to the dying neuro…

cs.LG2024

Efficient Search for Customized Activation Functions with Gradient Descent

Lukas Strack, Mahmoud Safari, Frank Hutter

Different activation functions work best for different deep learning models. To exploit this, we leverage recent advancements in gradient-based search techniques for neural archite…

cs.LG2024

Surprisingly Strong Performance Prediction with Neural Graph Features

Gabriela Kadlecová, Jovita Lukasik, Martin Pilát +4

Performance prediction has been a key part of the neural architecture search (NAS) process, allowing to speed up NAS algorithms by avoiding resource-consuming network training. Alt…