3 papers
cs.LG2026
Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision
Grzegorz Gruszczynski, Pawel Olszowiec, Michal Byra +2
Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is independently parameterized. Single-…
cs.AI2026
Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data
Grzegorz Stefanski, Alberto Presta, Michal Byra
In pruning, the Lottery Ticket Hypothesis posits that large networks contain sparse subnetworks, or winning tickets, that can be trained in isolation to match the performance of th…
cs.LG2024
SOI: Scaling Down Computational Complexity by Estimating Partial States of the Model
Grzegorz Stefański, Paweł Daniluk, Artur Szumaczuk +1
Consumer electronics used to follow the miniaturization trend described by Moore's Law. Despite increased processing power in Microcontroller Units (MCUs), MCUs used in the smalles…