3 papers
cs.AI2026
Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy
Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introd…
cs.HC2025
PTFA: An LLM-based Agent that Facilitates Online Consensus Building through Parallel Thinking
Wen Gu, Zhaoxing Li, Jan Buermann +5
Consensus building is inherently challenging due to the diverse opinions held by stakeholders. Effective facilitation is crucial to support the consensus building process and enabl…
cs.CL2025
Reinforced Language Models for Sequential Decision Making
Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
Large Language Models (LLMs) show potential as sequential decision-making agents, but their application is often limited due to a reliance on large, computationally expensive model…