13 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.CL2025★ 13 cited
Deliberative Alignment: Reasoning Enables Safer Language Models
Melody Y. Guan, Manas Joglekar, Eric Wallace +12
As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introdu…
cs.CL2024
Thesis proposal: Are We Losing Textual Diversity to Natural Language Processing?
Josef Jon
This thesis argues that the currently widely used Natural Language Processing algorithms possibly have various limitations related to the properties of the texts they handle and pr…