From the 1 of 8 linked papers with an AI index.
8 papers
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
Renlong Wu, Haoran Chen, Yuxiang Wei +3
The paper introduces 4DHumanDiff, a diffusion-based framework that directly generates 360-degree dynamic human models as 4D Gaussian Splatting representations from text prompts, el…
Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation
Zihao Wang, Yuxiang Wei, Xinpeng Zhou +5
Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on multimodal large language model…
DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing
Kailai Feng, Yuxiang Wei, Bo Chen +5
Diffusion models have made significant progress in both text-to-image (T2I) generation and text-guided image editing. However, these models are typically built with billions of par…
CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
Chonghuinan Wang, Zihan Chen, Yuxiang Wei +5
Instruction-based multimodal image manipulation has recently made rapid progress. However, existing evaluation methods lack a systematic and human-aligned framework for assessing m…
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
Minheng Ni, Yutao Fan, Zhengyuan Yang +6
Recent advances in large multimodal models (LMMs) have enabled instruction-based image editing, allowing users to modify visual content via natural language descriptions. However,…
EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models
Yufei Cai, Hu Han, Yuxiang Wei +2
The progress on generative models has led to significant advances on text-to-video (T2V) generation, yet the motion controllability of generated videos remains limited. Existing mo…