3 papers
cs.CV2025
STORK: Faster Diffusion And Flow Matching Sampling By Resolving Both Stiffness And Structure-Dependence
Zheng Tan, Weizhen Wang, Andrea L. Bertozzi +1
Diffusion models (DMs) and flow-matching models have demonstrated remarkable performance in image and video generation. However, such models require a significant number of functio…
cs.CV2025
Dreamland: Controllable World Creation with Simulator and Generative Models
Sicheng Mo, Ziyang Leng, Leon Liu +3
Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering thei…
cs.CV2025
Embodied Scene Understanding for Vision Language Models via MetaVQA
Weizhen Wang, Chenda Duan, Zhenghao Peng +2
Vision Language Models (VLMs) demonstrate significant potential as embodied AI agents for various mobility applications. However, a standardized, closed-loop benchmark for evaluati…