27 papers
Evidence-RL: Towards Evidence-intensive Visual Reasoning
Haojie Huang, Xinlei Yu, Chengming Xu +6
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware pos…
CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation
Yuheng Chen, Teng Hu, Yuji Wang +7
The fidelity and structural diversity of training datasets fundamentally determine the capabilities of video generation models. While commercial systems showremarkableabilitytogene…
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
Yuheng Chen, Teng Hu, Yuji Wang +3
Identity-preserving video generation (IPVG) aims to synthesize high-fidelity videos that follow text prompts while faithfully preserving a reference identity. Despite recent progre…
SCL: Towards Domain Generalization via Single-Temporal Multimodal Contrastive Learning for Remote Sensing Change Detection
Qiangang Du, Jinlong Peng, Xu Chen +4
In recent years, change detection and anomaly detection models based on CNN and transformer have achieved remarkable success across various datasets based on paired data. However,…
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
Yanjie Pan, Qingdong He, Zhengkai Jiang +9
Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle…
PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset
Haojun Chen, Haoyang He, Chengming Xu +11
Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imagin…