Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
DRAFT: Task Decoupled Latent Reasoning for Agent Safety
Lin Wang, Junfeng Fang, Dan Zhang +3
The advent of tool-using LLM agents shifts safety monitoring from output moderation to auditing long, noisy interaction trajectories, where risk-critical evidence is sparse-making…
cs.LG2025
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
Jia Liu, ChangYi He, YingQiao Lin +3
Recent advancements in Large Language Models have yielded significant improvements in complex reasoning tasks such as mathematics and programming. However, these models remain heav…
cs.LG2025
Closing the Gap between TD Learning and Supervised Learning with -Conditioned Maximization
Xing Lei, Zifeng Zhuang, Shentao Yang +6
Recently, supervised learning (SL) methodology has emerged as an effective approach for offline reinforcement learning (RL) due to their simplicity, stability, and efficiency. Howe…