2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2023
Efficient Representation of the Activation Space in Deep Neural Networks
Tanya Akumu, Celia Cintas, Girmaw Abebe Tadesse +3
The representations of the activation space of deep neural networks (DNNs) are widely utilized for tasks like natural language processing, anomaly detection and speech recognition.…
cs.LG2023★ 2 cited
Weakly Supervised Detection of Hallucinations in LLM Activations
Miriam Rateike, Celia Cintas, John Wamburu +2
We propose an auditing method to identify whether a large language model (LLM) encodes patterns such as hallucinations in its internal states, which may propagate to downstream tas…