7 papers
PureCC: Pure Learning for Text-to-Image Concept Customization
Zhichao Liao, Xiaole Xian, Qingyu Li +7
Existing concept customization methods have achieved remarkable outcomes in high-fidelity and multi-concept customization. However, they often neglect the influence on the original…
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
Yingjie Ma, Xun Lin, Yong Xu +2
Face anti-spoofing (FAS) has recently advanced in multimodal fusion, cross-domain generalization, and interpretability. With large language models and reinforcement learning (RL),…
SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
Qilang Ye, Yu Zhou, Lian He +10
Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action…
Distribution-Specific Learning for Joint Salient and Camouflaged Object Detection
Chao Hao, Zitong Yu, Xin Liu +5
Salient object detection (SOD) and camouflaged object detection (COD) are two closely related but distinct computer vision tasks. Although both are class-agnostic segmentation task…
TRRG: Towards Truthful Radiology Report Generation With Cross-modal Disease Clue Enhanced Large Language Model
Yuhao Wang, Chao Hao, Yawen Cui +4
The vision-language modeling capability of multi-modal large language models has attracted wide attention from the community. However, in medical domain, radiology report generatio…
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
Bohao Xing, Zitong Yu, Xin Liu +6
Facial expression recognition (FER) is an important research topic in emotional artificial intelligence. In recent decades, researchers have made remarkable progress. However, curr…