#alignment
8 papers match
OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment
Seonglae Cho, Adriano Koshiyama
The paper presents OptimismBench, a benchmark that measures directional optimism or pessimism in large language models' probability judgments by comparing paired success/failure fo…
Constitutional Midtraining: Content Presence Drives Alignment Gains
Desiree Cho, Cameron Tice, Bernie Hogan +4
The paper investigates inserting constitutionally‑derived content during midtraining of large language models to improve the durability of alignment, showing reduced blackmail tend…
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
Anton de la Fuente, Arthur Conmy
The paper investigates whether lessons learned from supervised fine-tuning (SFT) in alignment training, model organisms, and toy models can be transferred across these domains, dem…
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
BuÄra Alperen Uluırmak, Rifat Kurban
The paper surveys recent work on evaluating large language models (LLMs) for safety and introduces the EvalSafetyGap framework to compare evaluation and alignment failures, illustr…
AvAtar: Learning to Align via Active Optimal Transport
Qi Yu, Ruizhong Qiu, Zhichen Zeng +3
The paper introduces AvAtar, an active learning framework that selects informative supervision points to improve optimal transport‑based alignment by measuring each candidate's gra…
Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments
Mingyu Liu, Zeju Li, Jiuhe Shu +4
The paper shows that smooth robot demonstrations can miss critical alignment moments, and proposes slowing down and resampling key motion segments, plus a spatio‑temporal feature c…
Align AI to Dynamic Human-AI Workflows
Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley +4
The paper proposes moving from static, preference‑emulating AI alignment toward interactive, complementary alignment where human and AI behaviors co‑evolve over time, drawing on so…
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
The paper introduces a benchmark of 128 tool‑calling scenarios to study how safety‑aligned large language models may override deployment instructions in regulated settings, reveali…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.