edge-cloud inference 1multimodal large language models 1query-guided pruning 1training-free methods 1visual token pruning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference
Feng Yang, Xinrui Ju, Keyang Zhang +6
The paper introduces LAST, a training‑free method that uses the attention of the last query token to prune visual tokens on edge devices before sending them to a cloud multimodal L…
cs.CV2025
Sparse2Dense: A Keypoint-driven Generative Framework for Human Video Compression and Vertex Prediction
Bolin Chen, Ru-Ling Liao, Yan Ye +5
For bandwidth-constrained multimedia applications, simultaneously achieving ultra-low bitrate human video compression and accurate vertex prediction remains a critical challenge, a…