5 citations · 12 across the 6 of their papers we have counts for
6 papers
Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems
Nikhil Khani, Shuo Yang, Aniruddh Nath +9
Knowledge Distillation (KD) is a powerful approach for compressing a large model into a smaller, more efficient model, particularly beneficial for latency-sensitive applications li…
COMET: Learning Cardinality Constrained Mixture of Experts with Trees and Local Search
Shibal Ibrahim, Wenyu Chen, Hussein Hazimeh +3
The sparse Mixture-of-Experts (Sparse-MoE) framework efficiently scales up model capacity in various domains, such as natural language processing and vision. Sparse-MoEs select a s…
Fast as CHITA: Neural Network Pruning with Combinatorial Optimization
Riade Benbaki, Wenyu Chen, Xiang Meng +4
The sheer size of modern neural networks makes model serving a serious computational challenge. A popular class of compression techniques overcomes this challenge by pruning or spa…
STTAR: Surgical Tool Tracking using off-the-shelf Augmented Reality Head-Mounted Displays
Alejandro Martin-Gomez, Haowei Li, Tianyu Song +6
The use of Augmented Reality (AR) for navigation purposes has shown beneficial in assisting physicians during the performance of surgical procedures. These applications commonly re…
M2HF: Multi-level Multi-modal Hybrid Fusion for Text-Video Retrieval
Shuo Liu, Weize Quan, Ming Zhou +5
Videos contain multi-modal content, and exploring multi-level cross-modal interactions with natural language queries can provide great prominence to text-video retrieval task (TVR)…
Scalable Bayesian Inference for Detection and Deblending in Astronomical Images
Derek Hansen, Ismael Mendoza, Runjing Liu +4
We present a new probabilistic method for detecting, deblending, and cataloging astronomical sources called the Bayesian Light Source Separator (BLISS). BLISS is based on deep gene…