2 papers
cs.CV2026
PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving
Lin Huang, Yujuan Tan, Weisheng Li +4
We present the PACE, a framework for retrieval-augmented dialogue serving that formalizes Perceived Time-to-First-Response (PTFR) as a QoE objective and minimizes it under quality/…
cs.LG2026
: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution Awareness
Lang Xiong, Ning Liu, Ao Ren +6
Large language models (LLMs) face significant deployment challenges due to their massive computational demands. % While pruning offers a promising compression solution, existing me…