2 papers
cs.LG2026
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
Renjie Mao, Xiangxin Zhou, Lvfang Tao +7
Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing PPO-style trust-region mechanisms remain position-agnostic…
cs.AI2026
InA-Probe: Instruction-Aware Active Probing for Time Series Forecasting with LLMs
Peiliang Gong, Emadeldeen Eldele, Chenyu Liu +8
Large Language Models (LLMs) have recently demonstrated impressive potential for time series forecasting. However, existing methods predominantly rely on passive modality alignment…