27 citations · 27 across the 2 of their papers we have counts for
2 papers
cs.CL2023
IFAN: An Explainability-Focused Interaction Framework for Humans and NLP Models
Edoardo Mosca, Daryna Dementieva, Tohid Ebrahim Ajdari +4
Interpretability and human oversight are fundamental pillars of deploying complex NLP models into real-world applications. However, applying explainability and human-in-the-loop me…
cs.AI2022★ 27 cited
"That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks
Edoardo Mosca, Shreyash Agarwal, Javier Rando +1
Adversarial attacks are a major challenge faced by current machine learning research. These purposely crafted inputs fool even the most advanced models, precluding their deployment…