2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.AI2024
Evaluating Readability and Faithfulness of Concept-based Explanations
Meng Li, Haoran Jin, Ruixuan Huang +5
With the growing popularity of general-purpose Large Language Models (LLMs), comes a need for more global explanations of model behaviors. Concept-based explanations arise as a pro…
cs.CL2024★ 2 cited
Uncovering Safety Risks of Large Language Models through Concept Activation Vector
Zhihao Xu, Ruixuan Huang, Changyu Chen +1
Despite careful safety alignment, current large language models (LLMs) remain vulnerable to various attacks. To further unveil the safety risks of LLMs, we introduce a Safety Conce…