3 papers
cs.CV2026
SpatialCrafter: Single Image World Modeling with Generative 3D Proxies
Chuan Fang, Lingteng Qiu, Yixun Liang +8
Explorable image-to-scene generation is essential for applications in gaming, robotics, and virtual reality. Existing methods based on video diffusion model (VDM) commonly rely on…
cs.CV2026
R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models
Qiwen Gu, Bingjie Gao, Rui Chen +7
High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very…
cs.CV2026
Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification
Zibo Zhou, Zongsen Qiu, Rui Chen +3
Automated tea leaf disease classification supports precision agriculture, yet deploying accurate models on edge devices remains challenging under tight compute budgets. Self-superv…