5 papers
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving
Ferran Agullo, Joan Oliveras, Chen Wang +5
Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of a…
Regression Models Meet Foundation Models: A Hybrid-AI Approach to Practical Electricity Price Forecasting
Yunzhong Qiu, Binzhu Li, Hao Wei +5
Electricity market prices exhibit extreme volatility, nonlinearity, and non-stationarity, making accurate forecasting a significant challenge. While cutting-edge time series founda…
Boosting AI Reliability with an FSM-Driven Streaming Inference Pipeline: An Industrial Case
Yutian Zhang, Zhongyi Pei, Yi Mao +3
The widespread adoption of AI in industry is often hampered by its limited robustness when faced with scenarios absent from training data, leading to prediction bias and vulnerabil…
Adapt Data to Model: Adaptive Transformation Optimization for Domain-shared Time Series Foundation Models
Yunzhong Qiu, Zhiyao Cen, Zhongyi Pei +2
Large time series models (LTMs) have emerged as powerful tools for universal forecasting, yet they often struggle with the inherent diversity and nonstationarity of real-world time…
BTTackler: A Diagnosis-based Framework for Efficient Deep Learning Hyperparameter Optimization
Zhongyi Pei, Zhiyao Cen, Yipeng Huang +4
Hyperparameter optimization (HPO) is known to be costly in deep learning, especially when leveraging automated approaches. Most of the existing automated HPO methods are accuracy-b…