4 papers
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
Zijin Yin, Bing Li, Kongming Liang +4
Semantic segmentation takes pivotal roles in various applications such as autonomous driving and medical image analysis. When deploying segmentation models in practice, it is criti…
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
Xinran Wang, Muxi Diao, Yuanzhi Liu +4
Training text-to-image (T2I) models with detailed captions can significantly improve their generation quality. Existing methods often rely on simplistic metrics like caption length…
ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
Jiayi Gao, Zijin Yin, Changcheng Hua +5
The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have t…
Detailed Object Description with Controllable Dimensions
Xinran Wang, Haiwen Zhang, Baoteng Li +5
Object description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models(MLLM…