Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Partial Fusion of Neural Networks: Efficient Tradeoffs Between Ensembles and Weight Aggregation
Fabian Morelli, Stephan Eckstein
Ensembles of neural networks typically outperform individual networks but incur large computational costs, whereas weight aggregation produces less costly, yet also less accurate,…
cs.LG2026
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
Moritz Brösamle, Moritz Brösamle, Stephan Eckstein
Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used…