Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
NFTR: From Provable Mode-Averaging to Geodesic Subgoal Selection in Offline Goal-Conditioned RL
Erdemt Bao, Xing Lei, Jun Chen
Hierarchical Implicit Q-Learning (HIQL), an offline goal-conditioned RL method, selects subgoals by value-function advantages alone. This rule has two coupled failure modes. Optimi…
cs.LG2026
ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning
Chengcao Yang
We propose a paradigm shift toward open-ended curriculum self-play: rather than learning to answer on a fixed prompt set, a unified policy learns to question: generating verifiable…
cs.LG2024
APT: Architectural Planning and Text-to-Blueprint Construction Using Large Language Models for Open-World Agents
Jun Yu Chen, Tao Gao
We present APT, an advanced Large Language Model (LLM)-driven framework that enables autonomous agents to construct complex and creative structures within the Minecraft environment…