Showing cs.PFShow all
2 papers · 1 filter
cs.PF2026
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
Tuowei Wang, He Zhou, Chengru Song +2
Large vision-language models (VLMs) are enabling interactive video reasoning, giving rise to streaming long-video understanding. In this setting, frames arrive continuously, while…
cs.PF2026
Swarm: Co-Activation Aware KVCache Offloading Across Multiple SSDs
Tuowei Wang, Liyun Chu, Ruwen Fan +1
The key-value (KV) cache has become the dominant contributor to memory consumption in large language model (LLM) inference. Although offloading KVCache from GPU high-bandwidth memo…