4 papers
CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels
Jirong Li, Satoshi Ikehata, Shuhei Kurita +1
3D Gaussian Splatting (3DGS) enables photorealistic real-time novel view synthesis, yet placing a virtual camera to capture a desired frame remains largely manual. Existing languag…
Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation
Sakuya Ota, Qing Yu, Kent Fujiwara +2
Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temporal, especially in composit…
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
Masahiro Kada, Ryota Yoshihashi, Satoshi Ikehata +2
Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottleneck. Sparse Mixture of Expert…
PINO: Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized Groups
Sakuya Ota, Qing Yu, Kent Fujiwara +2
Generating realistic group interactions involving multiple characters remains challenging due to increasing complexity as group size expands. While existing conditional diffusion m…