15 papers
HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression
Yuxuan Yang, Feiyang Ren, Bowen Zeng +4
Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top- nucleus sampling) offer superior accuracy by dynamically fluctuating me…
LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version)
Ankai Hao, Ke Chen, Huan Li +1
Feature engineering remains a cornerstone of tabular data analysis, and Large Language Models (LLMs) have emerged as a promising paradigm for its automation, giving rise to LLM-pow…
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
Yihao Wang, Haoran Xu, Renjie Gu +10
The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, ex…
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
Yicheng Ji, Zhizhou Zhong, Jun Zhang +7
Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness, as exemplified by the Self…
Token Economics for LLM Agents: A Dual-View Study from Computing and Economics
Yuxi Chen, Junming Chen, Chenyu He +9
As LLM agents evolve, tokens have emerged as the core economic primitives of Agentic AI. However, their exponential consumption introduces severe computational, collaborative, and…
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects
Jun Zhang, Yicheng Ji, Feiyang Ren +7
Large Vision-Language Models (LVLMs) enable sophisticated reasoning over images and videos, yet their inference is hindered by a systemic efficiency barrier known as visual token d…