activity
20242026
collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models

Chuhan Zhang, Seiji Ito, Kenta Hoshino +2

World Action Models (WAMs) aim to control robots by stochastically generating visual futures and then decoding actions, but empirical observations indicate that the results can str…

cs.CV2026

CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels

Jirong Li, Satoshi Ikehata, Shuhei Kurita +1

3D Gaussian Splatting (3DGS) enables photorealistic real-time novel view synthesis, yet placing a virtual camera to capture a desired frame remains largely manual. Existing languag…

cs.CV2026

Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation

Sakuya Ota, Qing Yu, Kent Fujiwara +2

Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temporal, especially in composit…

cs.CV2026

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization

Ryota Yoshihashi, Masahiro Kada, Satoshi Ikehata +2

Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, detection, and segmentation, a…

cs.CV2026

Teacher-Guided Routing for Sparse Vision Mixture-of-Experts

Masahiro Kada, Ryota Yoshihashi, Satoshi Ikehata +2

Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottleneck. Sparse Mixture of Expert…

cs.CV2025

Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues

King-Man Tam, Satoshi Ikehata, Yuta Asano +2

Universal Photometric Stereo is a promising approach for recovering surface normals without strict lighting assumptions. However, it struggles when multi-illumination cues are unre…