3 papers
cs.CV2026
TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models
Chang Sun, Francesco Barbato, Matteo Caligiuri +1
Vision-Language Models (VLMs) exhibit strong zero-shot capabilities, making them an attractive solution for continual learning across diverse tasks. However, during continual adapt…
cs.CV2026
LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
Enhuai Liu, Yunke Wang, Yutong Wang +2
Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and dura…
cs.CV2026
SLAMFormer-: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing
Zhijian Fang, Weicheng Zheng, Yijun Yuan +7
We introduce the Infinite SLAM Transformer (SLAMFormer-), the first geometric transformer capable of supporting both long-range frontend and backend processing without an e…