3 papers
cs.CL2026
Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning
Jinlong Liu, Mohammed Bahja, Venelin Kovatchev +1
Evaluating and optimising authorial style in long-form story generation remains challenging because style is often assessed with ad hoc prompting and is frequently conflated with o…
cs.CY2025
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…
cs.CL2025
Finding Pareto Trade-offs in Fair and Accurate Detection of Toxic Speech
Soumyajit Gupta, Venelin Kovatchev, Anubrata Das +2
Optimizing NLP models for fairness poses many challenges. Lack of differentiable fairness measures prevents gradient-based loss training or requires surrogate losses that diverge f…