2 papers
cs.DC2024
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
Zihan Chang, Sheng Xiao, Shuibing He +3
Existing work only effective on a given number of GPUs, often neglecting the complexities involved in manually determining the specific types and quantities of GPUs needed, which c…
cs.CV2024
MixBCT: Towards Self-Adapting Backward-Compatible Training
Yu Liang, Yufeng Zhang, Shiliang Zhang +4
Backward-compatible training circumvents the need for expensive updates to the old gallery database when deploying an advanced new model in the retrieval system. Previous methods a…