8 papers
Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
Rujin Liang, Zhongpu Chen, Yuhao Lei +1
While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely c…
ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
Hankun Kang, Xin Miao, Jianhao Chen +5
Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social envi…
Beyond Static Snapshots: Dynamic Modeling and Forecasting of Group-Level Value Evolution with Large Language Models
Qiankun Pi, Guixin Su, Jinliang Li +5
Social simulation is critical for mining complex social dynamics and supporting data-driven decision making. LLM-based methods have emerged as powerful tools for this task by lever…
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
Chen Qian, Xinran Yu, Danyang Li +4
Visual token pruning is a promising approach for reducing the computational cost of vision-language models (VLMs), and existing methods often rely on early pruning decisions to imp…
Aligning VLM Assistants with Personalized Situated Cognition
Yongqi Li, Shen Zhou, Xiaohu Li +9
Vision-language models (VLMs) aligned with general human objectives, such as being harmless and hallucination-free, have become valuable assistants of humans in managing visual tas…
Toxicity Detection towards Adaptability to Changing Perturbations
Hankun Kang, Jianhao Chen, Yongqi Li +5
Toxicity detection is crucial for maintaining the peace of the society. While existing methods perform well on normal toxic contents or those generated by specific perturbation met…