3 papers
cs.AI2026
DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo +4
Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet existing compression methods often apply…
cs.AI2026
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
Anna Córdoba, Adam Puente Tercero, Nerea Angulo Hijo +4
Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retri…
cs.CV2026
Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation
Anna Córdoba, Adam Puente Tercero, Nerea Angulo Hijo +4
We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, visual style, and causal progress…