2 papers
cs.LG2025
PETRA: Parallel End-to-end Training with Reversible Architectures
Stéphane Rivaud, Louis Fournier, Thomas Pumir +3
Reversible architectures have been shown to be capable of performing on par with their non-reversible architectures, being applied in deep learning for memory savings and generativ…
cs.LG2025
Growth strategies for arbitrary DAG neural architectures
Stella Douka, Manon Verbockhaven, Théo Rudkiewicz +4
Deep learning has shown impressive results obtained at the cost of training huge neural networks. However, the larger the architecture, the higher the computational, financial, and…