1 citations · 1 across the 3 of their papers we have counts for
3 papers
B-DENSE: Branching For Dense Ensemble Network Supervision Efficiency
Cherish Puniani, Tushar Kumar, Arnav Bendre +2
Inspired by non-equilibrium thermodynamics, diffusion models have achieved state-of-the-art performance in generative modeling. However, their iterative sampling nature results in…
Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
Leyla Mirvakhabova, Babak Ehteshami Bejnordi, Gaurav Kumar +3
Upcycling pre-trained dense models into sparse Mixture-of-Experts (MoEs) efficiently increases model capacity but often suffers from poor expert specialization due to naive weight…
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
Babak Ehteshami Bejnordi, Gaurav Kumar, Amelie Royer +3
Jointly learning multiple tasks with a unified model can improve accuracy and data efficiency, but it faces the challenge of task interference, where optimizing one task objective…