3 papers
cs.CV2025
Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
Aybora Koksal, A. Aydin Alatan
Recent advances in large language and vision-language models have enabled strong reasoning capabilities, yet they remain impractical for specialized domains like remote sensing, wh…
cs.CV2024
Causal Transformer for Fusion and Pose Estimation in Deep Visual Inertial Odometry
Yunus Bilge Kurt, Ahmet Akman, A. Aydın Alatan
In recent years, transformer-based architectures become the de facto standard for sequence modeling in deep learning frameworks. Inspired by the successful examples, we propose a c…
cs.CV2024
XoFTR: Cross-modal Feature Matching Transformer
Ãnder TuzcuoÄlu, Aybora Köksal, BuÄra Sofu +2
We introduce, XoFTR, a cross-modal cross-view method for local feature matching between thermal infrared (TIR) and visible images. Unlike visible images, TIR images are less suscep…