2 papers
cs.LG2026
Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching
Qianli Ma, Zhiqing Tang, Hanshuai Cui +2
Disaggregated serving alleviates memory bottlenecks in Large Language Model (LLM) inference but creates a severe communication bottleneck: transmitting high-dimensional Key-Value (…
cs.CV2026
Models as Lego Builders: Assembling Malice from Benign Blocks via Semantic Blueprints
Chenxi Li, Xianggan Liu, Dake Shen +9
Despite the rapid progress of Large Vision-Language Models (LVLMs), the integration of visual modalities introduces new safety vulnerabilities that adversaries can exploit to elici…