activity
20212026
most citedHate-Alert@DravidianLangTech-EACL2021: Ensembling strategies for Transformer-based Offensive language Detection

14 citations · 33 across the 15 of their papers we have counts for

collaborators
Showing cs.CLShow all

17 papers · 1 filter

cs.CL2026

Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models

Punyajoy Saha, Sudipta Halder, Debjyoti Mondal +1

Safety alignment is critical for deploying large language models (LLMs) in real-world applications, yet most existing approaches rely on large human-annotated datasets and static r…

cs.CL2025

HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation

Naquee Rizwan, Seid Muhie Yimam, Daryna Dementieva +11

Despite regulations imposed by nations and social media platforms, e.g. (Government of India, 2021; European Parliament and Council of the European Union, 2022), inter alia, hatefu…

cs.CL2024★ 1 cited

CrowdCounter: A benchmark type-specific multi-target counterspeech dataset

Punyajoy Saha, Abhilash Datta, Abhik Jana +1

Counterspeech presents a viable alternative to banning or suspending users for hate speech while upholding freedom of expression. However, writing effective counterspeech is challe…

cs.CL2024

Demarked: A Strategy for Enhanced Abusive Speech Moderation through Counterspeech, Detoxification, and Message Management

Seid Muhie Yimam, Daryna Dementieva, Tim Fischer +8

Despite regulations imposed by nations and social media platforms, such as recent EU regulations targeting digital violence, abusive content persists as a significant challenge. Ex…

cs.CL2024

On Zero-Shot Counterspeech Generation by LLMs

Punyajoy Saha, Aalok Agrawal, Abhik Jana +2

With the emergence of numerous Large Language Models (LLM), the usage of such models in various Natural Language Processing (NLP) applications is increasing extensively. Counterspe…

cs.CL2024

InfFeed: Influence Functions as a Feedback to Improve the Performance of Subjective Tasks

Somnath Banerjee, Maulindu Sarkar, Punyajoy Saha +2

Recently, influence functions present an apparatus for achieving explainability for deep neural models by quantifying the perturbation of individual train instances that might impa…