◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

D. Majercak

3 papers hereh-index 5172 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1

Across the 1 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG1

identity via Semantic Scholar / OpenAlex

most citedPhi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

2 citations · 2 across the 2 of their papers we have counts for

collaborators

3 papers

cs.CL2025

The Bias is in the Details: An Assessment of Cognitive Bias in LLMs

R. Alexander Knipper, Charles S. Knipper, Kaiqi Zhang +3

As Large Language Models (LLMs) are increasingly embedded in real-world decision-making processes, it becomes crucial to examine the extent to which they exhibit cognitive biases.…

cs.LG2024

Steering Language Model Refusal with Sparse Autoencoders

Kyle O'Brien, David Majercak, Xavier Fernandes +7

Responsible deployment of language models requires mechanisms for refusing unsafe prompts while preserving model performance. While most approaches modify model weights through add…

cs.CL2024★ 2 cited

Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

Emman Haider, Daniel Perez-Becker, Thomas Portet +28

Recent innovations in language model training have demonstrated that it is possible to create highly performant models that are small enough to run on a smartphone. As these models…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.