Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning
Yosub Shin, Michael Buriek, Igor Molybog
Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead repres…
cs.CV2025
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
Yosub Shin, Igor Molybog
Video synchronization-aligning multiple video streams capturing the same event from different angles-is crucial for applications such as reality TV show production, sports analysis…