2 papers
cs.CV2026
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models
Yu Chen, Xiaohong Li, Xiaole Wang +3
In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply.…
cs.CL2026
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
Yu Chen, Runkai Chen, Sheng Yi +9
Long-term memory is a cornerstone of human intelligence. Enabling AI to process lifetime-scale information remains a long-standing pursuit in the field. Due to the constraints of f…