14 citations · 33 across the 15 of their papers we have counts for
17 papers · 1 filter
Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models
Punyajoy Saha, Sudipta Halder, Debjyoti Mondal +1
Safety alignment is critical for deploying large language models (LLMs) in real-world applications, yet most existing approaches rely on large human-annotated datasets and static r…
HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation
Naquee Rizwan, Seid Muhie Yimam, Daryna Dementieva +11
Despite regulations imposed by nations and social media platforms, e.g. (Government of India, 2021; European Parliament and Council of the European Union, 2022), inter alia, hatefu…
CrowdCounter: A benchmark type-specific multi-target counterspeech dataset
Punyajoy Saha, Abhilash Datta, Abhik Jana +1
Counterspeech presents a viable alternative to banning or suspending users for hate speech while upholding freedom of expression. However, writing effective counterspeech is challe…
Demarked: A Strategy for Enhanced Abusive Speech Moderation through Counterspeech, Detoxification, and Message Management
Seid Muhie Yimam, Daryna Dementieva, Tim Fischer +8
Despite regulations imposed by nations and social media platforms, such as recent EU regulations targeting digital violence, abusive content persists as a significant challenge. Ex…
On Zero-Shot Counterspeech Generation by LLMs
Punyajoy Saha, Aalok Agrawal, Abhik Jana +2
With the emergence of numerous Large Language Models (LLM), the usage of such models in various Natural Language Processing (NLP) applications is increasing extensively. Counterspe…
InfFeed: Influence Functions as a Feedback to Improve the Performance of Subjective Tasks
Somnath Banerjee, Maulindu Sarkar, Punyajoy Saha +2
Recently, influence functions present an apparatus for achieving explainability for deep neural models by quantifying the perturbation of individual train instances that might impa…