2 papers
cs.CV2026
What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
Ryota Yoshihashi, Masahiro Kada, Satoshi Ikehata +2
Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, detection, and segmentation, a…
cs.CV2026
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
Masahiro Kada, Ryota Yoshihashi, Satoshi Ikehata +2
Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottleneck. Sparse Mixture of Expert…