Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
DRAFT: Task Decoupled Latent Reasoning for Agent Safety
Lin Wang, Junfeng Fang, Dan Zhang +3
The advent of tool-using LLM agents shifts safety monitoring from output moderation to auditing long, noisy interaction trajectories, where risk-critical evidence is sparse-making…
cs.LG2025
Closing the Gap between TD Learning and Supervised Learning with -Conditioned Maximization
Xing Lei, Zifeng Zhuang, Shentao Yang +6
Recently, supervised learning (SL) methodology has emerged as an effective approach for offline reinforcement learning (RL) due to their simplicity, stability, and efficiency. Howe…
cs.LG2025
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
Jia Liu, ChangYi He, YingQiao Lin +3
Recent advancements in Large Language Models have yielded significant improvements in complex reasoning tasks such as mathematics and programming. However, these models remain heav…