7 papers
HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression
Yuxuan Yang, Feiyang Ren, Bowen Zeng +4
Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top- nucleus sampling) offer superior accuracy by dynamically fluctuating me…
Token Economics for LLM Agents: A Dual-View Study from Computing and Economics
Yuxi Chen, Junming Chen, Chenyu He +9
As LLM agents evolve, tokens have emerged as the core economic primitives of Agentic AI. However, their exponential consumption introduces severe computational, collaborative, and…
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects
Jun Zhang, Yicheng Ji, Feiyang Ren +7
Large Vision-Language Models (LVLMs) enable sophisticated reasoning over images and videos, yet their inference is hindered by a systemic efficiency barrier known as visual token d…
See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs
Yicheng Ji, Jun Zhang, Jinpeng Chen +4
Video Large Language Models (Video-LLMs) excel in video understanding but suffer from high inference latency during autoregressive generation. Speculative Decoding (SD) mitigates t…
CTkvr: KV Cache Retrieval for Long-Context LLMs via Centroid then Token Indexing
Kuan Lu, Shuhang Lin, Sai Wu +7
Large language models (LLMs) are increasingly applied in long-context scenarios such as multi-turn conversations. However, long contexts pose significant challenges for inference e…
Not All Data are Good Labels: On the Self-supervised Labeling for Time Series Forecasting
Yuxuan Yang, Dalin Zhang, Yuxuan Liang +3
Time Series Forecasting (TSF) is a crucial task in various domains, yet existing TSF models rely heavily on high-quality data and insufficiently exploit all available data. This pa…