3 papers
cs.RO2025
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
Zhenhan Yin, Xuanhan Wang, Jiahao Jiang +8
While leveraging abundant human videos and simulated robot data poses a scalable solution to the scarcity of real-world robot data, the generalization capability of existing vision…
cs.CV2025
LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
Yang Xiao, Gen Li, Kaiyuan Deng +5
Training-free acceleration has emerged as an advanced research area in video generation based on diffusion models. The redundancy of latents in diffusion model inference provides a…
cs.CV2025
GS2E: Gaussian Splatting is an Effective Data Generator for Event Stream Generation
Yuchen Li, Chaoran Feng, Zhenyu Tang +4
We introduce GS2E (Gaussian Splatting to Event), a large-scale synthetic event dataset for high-fidelity event vision tasks, captured from real-world sparse multi-view RGB images.…