4 papers
FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
Peiyuan Zhang, Xiangyu Zhao, Hongbo Liu +8
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on…
RISE-Video: Can Video Generators Decode Implicit World Rules?
Mingxin Liu, Shuran Ma, Shibei Meng +9
While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored fro…
Object Fidelity Diffusion for Remote Sensing Image Generation
Ziqi Ye, Shuran Ma, Jie Yang +5
High-precision controllable remote sensing image generation is both meaningful and challenging. Existing diffusion models often produce low-fidelity images due to their inability t…
Co-Training Vision Language Models for Remote Sensing Multi-task Learning
Qingyun Li, Shuran Ma, Junwei Luo +8
With Transformers achieving outstanding performance on individual remote sensing (RS) tasks, we are now approaching the realization of a unified model that excels across multiple t…