3 papers
cs.CV2026
RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar +2
Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned…
cs.RO2026
SABER: A Scalable Action-Based Embodied Dataset for Real-World VLA Adaptation
Narsimha Menga, Parikshit Sakurikar, Amirreza Rouhi +6
Robotic deployment in real-world environments depends on rich, domain-specific action data as much as on strong model architecture. General-purpose robot foundation models show mod…
cs.CV2026
PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
Amirreza Rouhi, Parikshit Sakurikar, Satya Sai Reddy +6
A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deploy…