#alignment

try —

8 papers match

cs.CL2026

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

Seonglae Cho, Adriano Koshiyama

The paper presents OptimismBench, a benchmark that measures directional optimism or pessimism in large language models' probability judgments by comparing paired success/failure fo…

#bias detection#language model evaluation#probability calibration#alignment
cs.CL2026

Constitutional Midtraining: Content Presence Drives Alignment Gains

Desiree Cho, Cameron Tice, Bernie Hogan +4

The paper investigates inserting constitutionally‑derived content during midtraining of large language models to improve the durability of alignment, showing reduced blackmail tend…

#alignment#midtraining#large language models#constitutional AI
cs.LG2026

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

Anton de la Fuente, Arthur Conmy

The paper investigates whether lessons learned from supervised fine-tuning (SFT) in alignment training, model organisms, and toy models can be transferred across these domains, dem…

#supervised fine-tuning#alignment#model organisms#toy models
cs.AI2026

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

Buğra Alperen Uluırmak, Rifat Kurban

The paper surveys recent work on evaluating large language models (LLMs) for safety and introduces the EvalSafetyGap framework to compare evaluation and alignment failures, illustr…

#large language models#evaluation#ai safety#benchmarking
cs.LG2026

AvAtar: Learning to Align via Active Optimal Transport

Qi Yu, Ruizhong Qiu, Zhichen Zeng +3

The paper introduces AvAtar, an active learning framework that selects informative supervision points to improve optimal transport‑based alignment by measuring each candidate's gra…

#optimal transport#active learning#alignment#gradient-based selection
cs.RO2026

Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments

Mingyu Liu, Zeju Li, Jiuhe Shu +4

The paper shows that smooth robot demonstrations can miss critical alignment moments, and proposes slowing down and resampling key motion segments, plus a spatio‑temporal feature c…

#imitation learning#manipulation#alignment#data augmentation
cs.AI2026

Align AI to Dynamic Human-AI Workflows

Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley +4

The paper proposes moving from static, preference‑emulating AI alignment toward interactive, complementary alignment where human and AI behaviors co‑evolve over time, drawing on so…

#human-ai interaction#alignment#dynamic preferences#collaborative AI
cs.SE2026

ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs

Aryan Keluskar, Amrita Bhattacharjee, Huan Liu

The paper introduces a benchmark of 128 tool‑calling scenarios to study how safety‑aligned large language models may override deployment instructions in regulated settings, reveali…

#alignment#tool-calling#large language models#safety

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.