12 papers
Weather-Conditioned Depth Anything
Zhaoming Xu, Chan-Wei Hu, Kuan-Ru Huang +4
Monocular depth estimation foundation models, such as the Depth Anything series, have achieved remarkable performance across diverse domains. However, they still suffer from critic…
Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
Xiangbo Gao, Siyuan Yang, Ping He +12
We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generation, and the update becomes visib…
4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation
Zihao Zhu, Kuan-Ru Huang, Zhaoming Xu +6
High-resolution datasets are essential for advancing super-resolution (SR) and text-to-image (T2I) diffusion research. However, current publicly available datasets lack both the na…
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
Yuheng Wu, Xiangbo Gao, Tianhao Chen +4
Interactive real-time autoregressive video generation is essential for applications such as content creation and world modeling, where visual content must adapt to dynamically evol…
Physics-Aware Video Instance Removal Benchmark
Zirui Li, Xinghao Chen, Lingyu Jiang +5
Video Instance Removal (VIR) requires removing target objects while maintaining background integrity and physical consistency, such as specular reflections and illumination interac…
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
Xiangbo Gao, Mingyang Wu, Siyuan Yang +4
While recent generative video models have achieved remarkable visual realism and are being explored as world models, true physical simulation requires mastering both space and time…