2 papers
cs.LG2026
From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
Kaiyuan Deng, Hangyu Zheng, Minghai Qing +11
Deploying models, especially large language models (LLMs), is becoming increasingly attractive to a broader user base, including those without specialized expertise. However, due t…
cs.CV2024
Exploring Token Pruning in Vision State Space Models
Zheng Zhan, Zhenglun Kong, Yifan Gong +8
State Space Models (SSMs) have the advantage of keeping linear computational complexity compared to attention modules in transformers, and have been applied to vision tasks as a ne…