From the 1 of 6 linked papers with an AI index.
6 papers
ReART: Reference-Guided Retrieval and Refinement for Emotion-Aware Art Generation
Qianqian Tang, Jiayi Gao, Ting Lei +1
Emotion-aware artistic image generation requires a model to satisfy semantic content, artistic style, and target emotion simultaneously. The key challenge is that artistic captions…
Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair
Jiayi Gao, Changcheng Hua, Jiaqi Tang +2
Identity-preserving video generation aims to synthesize videos that follow natural-language instructions while maintaining the visual identity of a given subject. Recent commercial…
Music-to-Dance Generation via Atomic Movements
Xinhao Cai, Yixuan Sun, Minghang Zheng +4
The paper proposes a framework that generates dance motions from music by first planning a sequence of interpretable atomic movements and then synthesizing smooth motion, improving…
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
Jiayi Gao, Changcheng Hua, Qingchao Chen +2
Identity-preserving text-to-video (IPT2V) generation creates videos faithful to both a reference subject image and a text prompt. While fine-tuning large pretrained video diffusion…
Learn 3D VQA Better with Active Selection and Reannotation
Shengli Zhou, Yang Liu, Feng Zheng
3D Visual Question Answering (3D VQA) is crucial for enabling models to perceive the physical world and perform spatial reasoning. In 3D VQA, the free-form nature of answers often…
ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
Jiayi Gao, Zijin Yin, Changcheng Hua +5
The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have t…