From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
Brady Bhalla, Honglu Fan, Nancy Chen +1
The paper studies how the size of embedding vectors influences the development of internal world models in transformers trained via reinforcement learning to perform bubble‑sort‑st…
cs.AI2026
The Reward Model Selection Crisis in Personalized Alignment
Fady Rezk, Yuangang Pan, Chuan-Sheng Foo +4
Personalized alignment from preference data has focused primarily on improving personal reward model (RM) accuracy, with the implicit assumption that better preference ranking tran…