2 papers
cs.LG2026
Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN
Tianhang Ding, Jianchun Liu, Hongli Xu
AI-RAN brings large language model (LLM) serving close to mobile users, but cellular handover can separate an active request from its inference state: the user attaches to a target…
cs.CV2026
SOS! : A Streamlined Object-Conditional Transformer for Model-free Segmentation
Jiaqi Hu, Junwen Huang, Hongli Xu +4
Foundation segmentation models excel at generating high-quality, class-agnostic masks, but they struggle to associate these proposals with specific target objects. This semantic ga…