1 citations · 1 across the 18 of their papers we have counts for
34 papers · 1 filter
Evidence-RL: Towards Evidence-intensive Visual Reasoning
Haojie Huang, Xinlei Yu, Chengming Xu +6
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware pos…
CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation
Yuheng Chen, Teng Hu, Yuji Wang +7
The fidelity and structural diversity of training datasets fundamentally determine the capabilities of video generation models. While commercial systems showremarkableabilitytogene…
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
Yuheng Chen, Teng Hu, Yuji Wang +3
Identity-preserving video generation (IPVG) aims to synthesize high-fidelity videos that follow text prompts while faithfully preserving a reference identity. Despite recent progre…
PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset
Haojun Chen, Haoyang He, Chengming Xu +11
Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imagin…
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
Yuheng Chen, Qingdong He, Teng Hu +4
The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides, achieving cohesive multimod…
CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal
Qingdong He, Chaoyi Wang, Peng Tang +2
Video subtitle removal aims to distinguish text overlays from background content while preserving temporal coherence. Existing diffusion-based methods necessitate explicit mask seq…