6 papers
Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation
Sakuya Ota, Qing Yu, Kent Fujiwara +2
Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temporal, especially in composit…
What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
Ryota Yoshihashi, Masahiro Kada, Satoshi Ikehata +2
Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, detection, and segmentation, a…
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
Masahiro Kada, Ryota Yoshihashi, Satoshi Ikehata +2
Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottleneck. Sparse Mixture of Expert…
Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues
King-Man Tam, Satoshi Ikehata, Yuta Asano +2
Universal Photometric Stereo is a promising approach for recovering surface normals without strict lighting assumptions. However, it struggles when multi-illumination cues are unre…
PINO: Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized Groups
Sakuya Ota, Qing Yu, Kent Fujiwara +2
Generating realistic group interactions involving multiple characters remains challenging due to increasing complexity as group size expands. While existing conditional diffusion m…
Rectified Lagrangian for Out-of-Distribution Detection in Modern Hopfield Networks
Ryo Moriai, Nakamasa Inoue, Masayuki Tanaka +3
Modern Hopfield networks (MHNs) have recently gained significant attention in the field of artificial intelligence because they can store and retrieve a large set of patterns with…