activity
20242026
collaborators

5 papers

cs.CV2026

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization

Ryota Yoshihashi, Masahiro Kada, Satoshi Ikehata +2

Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, detection, and segmentation, a…

cs.CV2026

Teacher-Guided Routing for Sparse Vision Mixture-of-Experts

Masahiro Kada, Ryota Yoshihashi, Satoshi Ikehata +2

Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottleneck. Sparse Mixture of Expert…

cs.CV2026

Constant Rate Scheduling: A General Framework for Optimizing Diffusion Noise Schedule via Distributional Change

Shuntaro Okada, Kenji Doi, Ryota Yoshihashi +2

We propose a general framework for optimizing noise schedules in diffusion models, applicable to both training and sampling. Our method enforces a constant rate of change in the pr…

cs.CV2025

VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction

Jiahao Zhang, Ryota Yoshihashi, Shunsuke Kitada +2

Large language models (LLMs) have proven effective for layout generation due to their ability to produce structure-description languages, such as HTML or JSON. In this paper, we ar…

cs.CV2024

Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation

Ryota Yoshihashi, Yuya Otsuka, Kenji Doi +2

The advance of generative models for images has inspired various training techniques for image recognition utilizing synthetic images. In semantic segmentation, one promising appro…