Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
Christian Simon, Masato Ishii, Wei-Yao Wang +8
Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information.…
cs.CV2026
SF-Mamba: Rethinking State Space Model for Vision
Masakazu Yoshimura, Teruaki Hayashi, Yuki Hoshino +2
The realm of Mamba for vision has been advanced in recent years to strike for the alternatives of Vision Transformers (ViTs) that suffer from the quadratic complexity. While the re…