2 citations · 2 across the 3 of their papers we have counts for
3 papers
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
The AIBrix Team, Jiaxin Shan, Varun Gupta +24
We introduce AIBrix, a cloud-native, open-source framework designed to optimize and simplify large-scale LLM deployment in cloud environments. Unlike traditional cloud-native stack…
GPU Memory Usage Optimization for Backward Propagation in Deep Network Training
Ding-Yong Hong, Tzu-Hsien Tsai, Ning Wang +2
In modern Deep Learning, it has been a trend to design larger Deep Neural Networks (DNNs) for the execution of more complex tasks and better accuracy. On the other hand, Convolutio…
Strategies for Optimizing End-to-End Artificial Intelligence Pipelines on Intel Xeon Processors
Meena Arunachalam, Vrushabh Sanghavi, Yi A Yao +6
End-to-end (E2E) artificial intelligence (AI) pipelines are composed of several stages including data preprocessing, data ingestion, defining and training the model, hyperparameter…