3 papers
cs.AI2026
Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy
Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introd…
cs.CL2025
Reinforced Language Models for Sequential Decision Making
Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
Large Language Models (LLMs) show potential as sequential decision-making agents, but their application is often limited due to a reliance on large, computationally expensive model…
cs.HC2025
PTFA: An LLM-based Agent that Facilitates Online Consensus Building through Parallel Thinking
Wen Gu, Zhaoxing Li, Jan Buermann +5
Consensus building is inherently challenging due to the diverse opinions held by stakeholders. Effective facilitation is crucial to support the consensus building process and enabl…