3 papers
cs.CL2026
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
Zhiyuan Shi, Qibo Qiu, Feng Xue +5
The linear memory growth of the KV cache poses a significant bottleneck for LLM inference in long-context tasks. Existing static compression methods often fail to preserve globally…
cs.RO2026
EdgeNav-QE: QLoRA Quantization and Dynamic Early Exit for LAM-based Navigation on Edge Devices
Mengyun Liu, Shanshan Huang, Jianan Jiang
Large Action Models (LAMs) have shown immense potential in autonomous navigation by bridging high-level reasoning with low-level control. However, deploying these multi-billion par…
cs.CV2024
CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection
Qibo Chen, Weizhong Jin, Jianyue Ge +7
Recent research on universal object detection aims to introduce language in a SoTA closed-set detector and then generalize the open-set concepts by constructing large-scale (text-r…