9 papers
Text-Vision Co-Instructed Image Editing
Chenxi Xie, Yuhui Wu, Qiaosi Yi +1
Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are l…
CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning
Yuhui Wu, Chenxi Xie, Ruibin Li +3
Image editing has achieved impressive results with the development of large-scale generative models. However, existing models mainly focus on the editing effects of intended object…
T2M Mamba: Motion Periodicity-Saliency Coupling Approach for Stable Text-Driven Motion Generation
Xingzu Zhan, Chen Xie, Honghang Chen +2
Text-to-motion generation, which converts motion language descriptions into coherent 3D human motion sequences, has attracted increasing attention in fields, such as avatar animati…
FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models
Minghan Li, Chenxi Xie, Yichen Wu +2
Numerous text-to-video (T2V) editing methods have emerged recently, but the lack of a standardized benchmark for fair evaluation has led to inconsistent claims and an inability to…
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction
Yuhui Wu, Liyi Chen, Ruibin Li +3
Instruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting hig…
MaSS13K: A Matting-level Semantic Segmentation Benchmark
Chenxi Xie, Minghan Li, Hui Zeng +2
High-resolution semantic segmentation is essential for applications such as image editing, bokeh imaging, AR/VR, etc. Unfortunately, existing datasets often have limited resolution…