◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Tomas Hrycej

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG3

identity via Semantic Scholar / OpenAlex

most citedEfficient Neural Network Training via Subset Pretraining

3 citations · 4 across the 3 of their papers we have counts for

collaborators

3 papers

cs.LG2026

Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder

Götz-Henrik Wiegand, Lorena Raichle, Rico Städeli +3

Training Transformer language models is expensive, as performance typically improves with increasing dataset size and computational budget. Although scaling laws describe this tren…

cs.LG2024★ 3 cited

Efficient Neural Network Training via Subset Pretraining

Jan Spörer, Bernhard Bermeitinger, Tomas Hrycej +2

In training neural networks, it is common practice to use partial gradients computed over batches, mostly very small subsets of the training set. This approach is motivated by the…

cs.LG2024★ 1 cited

Reducing the Transformer Architecture to a Minimum

Bernhard Bermeitinger, Tomas Hrycej, Massimo Pavone +2

Transformers are a widespread and successful model architecture, particularly in Natural Language Processing (NLP) and Computer Vision (CV). The essential innovation of this archit…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.