2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2024
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
Yiping Wang, Xuehai He, Kuan Wang +5
The current state-of-the-art video generative models can produce commercial-grade videos with highly realistic details. However, they still struggle to coherently present multiple…
cs.CV2024
Mojito: Motion Trajectory and Intensity Control for Video Generation
Xuehai He, Shuohang Wang, Jianwei Yang +7
Recent advancements in diffusion models have shown great promise in producing high-quality video content. However, efficiently training video diffusion models capable of integratin…
cs.CV2024★ 2 cited
OmniParser for Pure Vision Based GUI Agent
Yadong Lu, Jianwei Yang, Yelong Shen +1
The recent success of large vision language models shows great potential in driving the agent system operating on user interfaces. However, we argue that the power multimodal model…