6 papers
PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving
Meghana Maghyastha, Robert Underwood, Randal Burns +1
Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the…
Recency/Frequency Adaptive KV Caching for Large Language Model Serving
Yang Shen, Meghana Madhyastha, Robert Underwood +2
Key-value (KV) caching is a powerful technique for accelerating large language model inference and generation. Inference workloads are large and diverse, which makes them difficult…
On Harnessing Idle Compute at the Edge for Foundation Model Training
Leyang Xue, Meghana Madhyastha, Myungjin Lee +3
The foundation-model ecosystem remains highly centralized because training requires immense compute resources and is therefore largely limited to large cloud operators. Edge-assist…
Vectorized Adaptive Histograms for Sparse Oblique Forests
Ariel Lubonja, Jungsang Yoon, Haoyin Xu +6
Classification using sparse oblique random forests provides guarantees on uncertainty and confidence while controlling for specific error types. However, they use more data and mor…
Auditing Significance, Metric Choice, and Demographic Fairness in Medical AI Challenges
Ariel Lubonja, Pedro R. A. S. Bassi, Wenxuan Li +4
Open challenges have become the de facto standard for comparative ranking of medical AI methods. Despite their importance, medical AI leaderboards exhibit three persistent limitati…
Towards Decentralized and Sustainable Foundation Model Training with the Edge
Leyang Xue, Meghana Madhyastha, Randal Burns +2
Foundation models are at the forefront of AI research, appealing for their ability to learn from vast datasets and cater to diverse tasks. Yet, their significant computational dema…