large language models 1post-training dynamics 1preference optimization 1supervised fine-tuning 1value alignment 1
From the 1 of 14 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Operationalising the Superficial Alignment Hypothesis via Task Complexity
Tomás Vergara-Browne, Darshan Patil, Ivan Titov +3
The superficial alignment hypothesis (SAH) posits that large language models learn most of their knowledge during pre-training, and that post-training merely surfaces this knowledg…
cs.LG2025
Build the web for agents, not agents for the web
Xing Han Lù, Gaurav Kamath, Marius Mosbach +1
Recent advancements in Large Language Models (LLMs) and multimodal counterparts have spurred significant interest in developing web agents -- AI systems capable of autonomously nav…