◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

David D. Baek

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.AI1

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2025

Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth

Jiawei Zhang, Andrew Estornell, David D. Baek +2

Large Language Models (LLMs) exhibit strong but shallow alignment: they directly refuse harmful queries when a refusal is expected at the very start of an assistant turn, yet this…

cs.AI2025

Scaling Laws For Scalable Oversight

Joshua Engels, David D. Baek, Subhash Kantamneni +1

Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems. However, it is s…

cs.LG2025

Towards Understanding Distilled Reasoning Models: A Representational Approach

David D. Baek, Max Tegmark

In this paper, we investigate how model distillation impacts the development of reasoning features in large language models (LLMs). To explore this, we train a crosscoder on Qwen-s…

cs.LG2025

Harmonic Loss Trains Interpretable AI Models

David D. Baek, Ziming Liu, Riya Tyagi +1

In this paper, we introduce harmonic loss as an alternative supervisory signal for training neural networks and large language models (LLMs). Harmonic loss differs from standard cr…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.