Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed un…
cs.CL2023
Developing Linguistic Patterns to Mitigate Inherent Human Bias in Offensive Language Detection
Toygar Tanyel, Besher Alkurdi, Serkan Ayvaz
With the proliferation of social media, there has been a sharp increase in offensive content, particularly targeting vulnerable groups, exacerbating social problems such as hatred,…