1 citations · 1 across the 2 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
REFINE-LM: Mitigating Language Model Stereotypes via Reinforcement Learning
Rameez Qureshi, Naïm Es-Sebbani, Luis Galárraga +3
With the introduction of (large) language models, there has been significant concern about the unintended bias such models may inherit from their training data. A number of studies…
cs.CL2024★ 1 cited
Does It Make Sense to Explain a Black Box With Another Black Box?
Julien Delaunay, Luis Galárraga, Christine Largouët
Although counterfactual explanations are a popular approach to explain ML black-box classifiers, they are less widespread in NLP. Most methods find those explanations by iterativel…