Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar +2
Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned…
cs.CV2026
PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
Amirreza Rouhi, Parikshit Sakurikar, Satya Sai Reddy +6
A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deploy…