Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
Zhe Li, Wei Zhao, Yige Li +1
Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their deployment is frequently undermined by undesirable behaviors such as generating harmful content, f…
cs.CL2025
Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs
Wei Zhao, Zhe Li, Yige Li +1
Large Vision-Language Models (LVLMs) have made significant strides in multimodal comprehension, thanks to extensive pre-training and fine-tuning on large-scale visual datasets. How…
cs.CL2024
Do Influence Functions Work on Large Language Models?
Zhe Li, Wei Zhao, Yige Li +1
Influence functions are important for quantifying the impact of individual training data points on a model's predictions. Although extensive research has been conducted on influenc…