◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yassir Akram

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG3

identity via Semantic Scholar / OpenAlex

most citedRandom initialisations performing above chance and how to find them

3 citations · 3 across the 3 of their papers we have counts for

collaborators

3 papers

cs.LG2024

Weight decay induces low-rank attention layers

Seijin Kobayashi, Yassir Akram, Johannes Von Oswald

The effect of regularizers such as weight decay when training deep neural networks is not well understood. We study the influence of weight decay as well as L2-regularization whe…

cs.LG2024

Learning Randomized Algorithms with Transformers

Johannes von Oswald, Seijin Kobayashi, Yassir Akram +1

Randomization is a powerful tool that endows algorithms with remarkable properties. For instance, randomized algorithms excel in adversarial settings, often surpassing the worst-ca…

cs.LG2022★ 3 cited

Random initialisations performing above chance and how to find them

Frederik Benzing, Simon Schug, Robert Meier +5

Neural networks trained with stochastic gradient descent (SGD) starting from different random initialisations typically find functionally very similar solutions, raising the questi…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.