7 papers
Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
Jinhao Pan, Chahat Raj, Anjishnu Mukherjee +4
Large language models (LLMs) exhibit social biases that reinforce harmful stereotypes, limiting their safe deployment. Most existing debiasing methods adopt a suppressive paradigm…
Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback
Bowen Wei, Nan Wang, Yuqing Zhou +2
Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. However, in many reasoning do…
A Logical-Rule Autoencoder for Interpretable Recommendations
Jinhao Pan, Bowen Wei, Ziwei Zhu
Most deep learning recommendation models operate as black boxes, relying on latent representations that obscure their decision process. This lack of intrinsic interpretability rais…
ClawSafety: "Safe" LLMs, Unsafe Agents
Bowen Wei, Yunbei Zhang, Jinhao Pan +5
Personal AI agents like OpenClaw run with elevated privileges on users' local machines, where a single successful prompt injection can leak credentials, redirect financial transact…
Bias Association Discovery Framework for Open-Ended LLM Generations
Jinhao Pan, Chahat Raj, Ziwei Zhu
Social biases embedded in Large Language Models (LLMs) raise critical concerns, resulting in representational harms -- unfair or distorted portrayals of demographic groups -- that…
CORTEX: Collaborative LLM Agents for High-Stakes Alert Triage
Bowen Wei, Yuan Shen Tay, Howard Liu +4
Security Operations Centers (SOCs) are overwhelmed by tens of thousands of daily alerts, with only a small fraction corresponding to genuine attacks. This overload creates alert fa…