3 papers
cs.LG2026
Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts
Dohyeon Kim, Bedionita Soro, Sung Ju Hwang
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most…
cs.LG2025
Instruction-Guided Autoregressive Neural Network Parameter Generation
Soro Bedionita, Bruno Andreis, Song Chong +1
Learning to generate neural network parameters conditioned on task descriptions and architecture specifications is pivotal for advancing model adaptability and transfer learning. E…
cs.LG2024
Diffusion-Based Neural Network Weights Generation
Bedionita Soro, Bruno Andreis, Hayeon Lee +4
Transfer learning has gained significant attention in recent deep learning research due to its ability to accelerate convergence and enhance performance on new tasks. However, its…