evaluation metrics 1model probing 1preference modeling 1relational encoding 1RLHF 1transformer internals 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
Jan Kirin
Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B l…
cs.LG2026
Relational Preference Encoding in Looped Transformer Internal States
Jan Kirin
The paper studies how looped transformer models represent human preferences by training small evaluator heads on frozen model states, and after correcting evaluation errors it find…