2 papers
cs.LG2025
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
Ashwinee Panda, Vatsal Baherwani, Zain Sarwar +4
Mixture of Experts (MoE) pretraining is more scalable than dense Transformer pretraining, because MoEs learn to route inputs to a sparse set of their feedforward parameters. Howeve…
cs.CV2022
Out-of-Distribution Detection for LiDAR-based 3D Object Detection
Chengjie Huang, Van Duong Nguyen, Vahdat Abdelzad +5
3D object detection is an essential part of automated driving, and deep neural networks (DNNs) have achieved state-of-the-art performance for this task. However, deep models are no…