2 papers
cs.AI2025
Accelerating Multi-modal LLM Gaming Performance via Input Prediction and Mishit Correction
Ziyang Lin, Zixuan Sun, Sanhorn Chen +2
Real-time sequential control agents are often bottlenecked by inference latency. Even modest per-step planning delays can destabilize control and degrade overall performance. We pr…
cs.DC2025
Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
Jianli Jin, Ziyang Lin, Qianli Dong +5
With the proliferation of edge AI applications, satisfying user quality of experience (QoE) requirements, such as model inference latency, has become a first class objective, as th…