4 papers
HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes
Yujia Li, Yiqun Zhang, Zihan Cheng +7
Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfulness while misidentifying the attacked tar…
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
YiJie Huang, Yiqun Zhang, Zhuoyue Jia +7
Vision-language models improve perception by feeding increasingly long visual token sequences into language backbones, but the resulting inference cost raises a basic scaling quest…
Minstrel: Structural Prompt Generation with Multi-Agents Coordination for Non-AI Experts
Ming Wang, Yuanzhong Liu, Xiaoyu Liang +8
LLMs have demonstrated commendable performance across diverse domains. Nevertheless, formulating high-quality prompts to assist them in their work poses a challenge for non-AI expe…
Affective Computing in the Era of Large Language Models: A Survey from the NLP Perspective
Yiqun Zhang, Xiaocui Yang, Xingle Xu +8
Affective Computing (AC) integrates computer science, psychology, and cognitive science to enable machines to recognize, interpret, and simulate human emotions across domains such…