From the 1 of 4 linked papers with an AI index.
4 papers
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Manas Pathak, Xingyao Chen, Shuozhe Li +2
The paper introduces the Filtered Reasoning Score (FRS), a metric that evaluates the quality of reasoning traces from large language models by focusing on the most confident genera…
Reinforcement Learning via Value Gradient Flow
Haoran Xu, Kaiwen Hu, Somayeh Sojoudi +1
We study behavior-regularized reinforcement learning (RL), where regularization toward a reference distribution (the dataset in offline RL or the base model in LLM RL finetuning) i…
AI and the Future of Digital Public Squares
Beth Goldberg, Diana Acosta-Navas, Michiel Bakker +24
Two substantial technological advances have reshaped the public square in recent decades: first with the advent of the internet and second with the recent introduction of large lan…
Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment
Andrew Konya, Aviv Ovadya, Kevin Feng +4
We introduce a method to measure the alignment between public will and language model (LM) behavior that can be applied to fine-tuning, online oversight, and pre-release safety che…