2 papers
cs.CV2025
A Study of Finetuning Video Transformers for Multi-view Geometry Tasks
Huimin Wu, Kwang-Ting Cheng, Stephen Lin +1
This paper presents an investigation of vision transformer learning for multi-view geometry tasks, such as optical flow estimation, by fine-tuning video foundation models. Unlike p…
cs.LG2025
Associative Transformer
Yuwei Sun, Hideya Ochiai, Zhirong Wu +2
Emerging from the pairwise attention in conventional Transformers, there is a growing interest in sparse attention mechanisms that align more closely with localized, contextual lea…