10 papers
Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
Jinhao Pan, Chahat Raj, Anjishnu Mukherjee +4
Large language models (LLMs) exhibit social biases that reinforce harmful stereotypes, limiting their safe deployment. Most existing debiasing methods adopt a suppressive paradigm…
Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback
Bowen Wei, Nan Wang, Yuqing Zhou +2
Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. However, in many reasoning do…
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models
Chahat Raj, Bowen Wei, Aylin Caliskan +2
While bias in large language models (LLMs) is well-studied, similar concerns in vision-language models (VLMs) have received comparatively less attention. Existing VLM bias studies…
A Logical-Rule Autoencoder for Interpretable Recommendations
Jinhao Pan, Bowen Wei, Ziwei Zhu
Most deep learning recommendation models operate as black boxes, relying on latent representations that obscure their decision process. This lack of intrinsic interpretability rais…
ClawSafety: "Safe" LLMs, Unsafe Agents
Bowen Wei, Yunbei Zhang, Jinhao Pan +5
Personal AI agents like OpenClaw run with elevated privileges on users' local machines, where a single successful prompt injection can leak credentials, redirect financial transact…
Interpretable Classification via a Rule Network with Selective Logical Operators
Bowen Wei, Ziwei Zhu
We introduce the Rule Network with Selective Logical Operators (RNS), a novel neural architecture that employs \textbf{selective logical operators} to adaptively choose between AND…