16 papers
Learning Explicit Physical Parameter Control and Benchmarking for Video Generation
Yanxun Li, Hao Wen, Bingze Song +7
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Cu…
OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data
Kaixing Yang, Jiashu Zhu, Xulong Tang +8
Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress…
GeoVista: Visually Grounded Active Perception for Vision-Language Understanding of Ultra-High-Resolution Remote Sensing Images
Jiashun Zhu, Ronghao Fu, Jiasen Hu +3
Interpreting ultra-high-resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large-scale scenes. Existing remote sensing visi…
DreamX-World 1.0: A General-Purpose Interactive World Model
DreamX Team, Yancheng Bai, Rui Chen +20
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…
SkyNative: A Native Multimodal Architecture for Remote Sensing Vision-Language Understanding
Xiao Yang, Ronghao Fu, Zhiwen Lin +10
Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the token space of a large language m…
Embedding-perturbed Exploration Preference Optimization for Flow Models
Sujie Hu, Chubin Chen, Jiashu Zhu +3
Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-based optimization framework…