3 papers
cs.CV2026
Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
Manyi Yao, Abhishek Aich, Yumin Suh +3
Vision transformer based models bring significant improvements for image segmentation tasks. Although these architectures offer powerful capabilities irrespective of specific segme…
cs.CV2025
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
Abhishek Aich, Yumin Suh, Samuel Schulter +1
A powerful architecture for universal segmentation relies on transformers that encode multi-scale image features and decode object queries into mask predictions. With efficiency be…
cs.CV2025
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
Dohwan Ko, Sihyeon Kim, Yumin Suh +4
Spatio-temporal reasoning is essential in understanding real-world environments in various fields, eg, autonomous driving and sports analytics. Recent advances have improved the sp…