1 citations · 2 across the 8 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
DMRL: Data- and Model-aware Reward Learning for Data Extraction
Zhiqiang Wang, Ruoxi Cheng
Large language models (LLMs) are inherently vulnerable to unintended privacy breaches. Consequently, systematic red-teaming research is essential for developing robust defense mech…
cs.LG2024
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
Shuirong Cao, Ruoxi Cheng, Zhiqiang Wang
LLMs can exhibit age biases, resulting in unequal treatment of individuals across age groups. While much research has addressed racial and gender biases, age bias remains little ex…
cs.LG2024
TUNI: A Textual Unimodal Detector for Identity Inference in CLIP Models
Songze Li, Ruoxi Cheng, Xiaojun Jia
The widespread usage of large-scale multimodal models like CLIP has heightened concerns about the leakage of PII. Existing methods for identity inference in CLIP models require que…