3 papers
cs.AI2026
KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning
Kun Feng, Ziwei Shan, Yuchen Fang +6
Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain semantic understanding, and effec…
cs.LG2026
Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
Zheng Zhang, Ziwei Shan, Kaitao Song +2
Process Reward Models (PRMs) have emerged as a promising approach to enhance the reasoning capabilities of large language models (LLMs) by guiding their step-by-step reasoning towa…
cs.LG2026
Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning
Zheng Zhang, Ao Lu, Yuanhao Zeng +5
Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant breakthroughs in complex LLM reasoning within verifiable domains, such as mathematics and programmin…