works on

From the 1 of 17 linked papers with an AI index.

collaborators

17 papers

cs.RO2026

Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

Yuzhou Wu, Yuxin Zheng, Muchun Niu +6

The paper introduces a system-level acceleration for vision-language-action models by incrementally updating visual tokens for dynamic regions and compressing diffusion-based polic…

cs.CL2026

Large Language Models Do Not Always Need Readable Language

Jiayi Zhu, Haoxuan Peng, Junxi Wang +3

Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model. This paper investigates whet…

cs.LG2026

dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching

Zhiyuan Liu, Yicun Yang, Yaojie Zhang +6

Autoregressive Models (ARMs) have long dominated the landscape of Large Language Models. Recently, a new paradigm has emerged in the form of diffusion-based Large Language Models (…

cs.CV2026

Where and How to Prune: An Empirical Study of Visual Token Pruning for GUI Agent Navigation

Daiqiang Li, Zihao Pan, Zeyu Zhang +8

In recent years, GUI agents have demonstrated strong potential in navigation tasks. However, preserving complete historical screenshots introduces substantial computational overhea…

cs.CV2026

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding

Junxi Wang, Te Sun, Jiayi Zhu +6

Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memory overhead, leading to high…

cs.SD2026

AudioKV: KV Cache Eviction in Efficient Large Audio Language Models

Yuxuan Wang, Peize He, Xiyan Gui +6

Large Audio-Language Models (LALMs) have set new benchmarks in speech processing, yet their deployment is hindered by the memory footprint of the Key-Value (KV) cache during long-c…