5 papers
What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
Ryota Yoshihashi, Masahiro Kada, Satoshi Ikehata +2
Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, detection, and segmentation, a…
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
Masahiro Kada, Ryota Yoshihashi, Satoshi Ikehata +2
Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottleneck. Sparse Mixture of Expert…
Constant Rate Scheduling: A General Framework for Optimizing Diffusion Noise Schedule via Distributional Change
Shuntaro Okada, Kenji Doi, Ryota Yoshihashi +2
We propose a general framework for optimizing noise schedules in diffusion models, applicable to both training and sampling. Our method enforces a constant rate of change in the pr…
VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction
Jiahao Zhang, Ryota Yoshihashi, Shunsuke Kitada +2
Large language models (LLMs) have proven effective for layout generation due to their ability to produce structure-description languages, such as HTML or JSON. In this paper, we ar…
Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation
Ryota Yoshihashi, Yuya Otsuka, Kenji Doi +2
The advance of generative models for images has inspired various training techniques for image recognition utilizing synthetic images. In semantic segmentation, one promising appro…