Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
CALM: Curiosity-Driven Auditing for Large Language Models
Xiang Zheng, Longxiang Wang, Yi Liu +3
Auditing Large Language Models (LLMs) is a crucial and challenging task. In this study, we focus on auditing black-box LLMs without access to their parameters, only to the provided…
cs.AI2024
Constrained Intrinsic Motivation for Reinforcement Learning
Xiang Zheng, Xingjun Ma, Chao Shen +1
This paper investigates two fundamental problems that arise when utilizing Intrinsic Motivation (IM) for reinforcement learning in Reward-Free Pre-Training (RFPT) tasks and Explora…