agentic retrieval 1dynamic knowledge graphs 1graph-based reasoning 1markov decision process 1multimodal reinforcement learning 1multimodal VQA 1policy alignment 1retrieval-augmented generation 1reward shaping 1visual grounding 1visual intervention 1
From the 2 of 8 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions
Jianghan Shen, Siqi Luo, Yue Li +9
Policy gradient algorithms for language models optimize the same objective , which has exactly two factors: the trajectory probability…
cs.AI2026
Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models
Siqi Luo, Jianghan Shen, Yi Xin +9
Diffusion Multi-Modal Large Language Models (dMLLMs) are powerful for image generation, but optimizing them through reinforcement learning (RL) remains a major challenge. One prima…
cs.AI2026
CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG
Jianghan Shen, Siqi Luo, Xinyu Cheng +6
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only superv…