#reinforcement learning
195 resultsKalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning
Haodong Zhu, Yangyang Ren, Yanjing Li +4
The paper introduces Kalman-Guided Prompt Selection (KGPS), a method that treats prompt difficulty as a dynamic state estimated with a Kalman filter to adaptively choose prompts du…
Hierarchical Latent Reasoning for LLM-based Recommendation
Peiyu Hu, Siying Gu, Weihai Lu +8
The paper introduces HiLaR, a framework that uses hierarchical latent reasoning and layer-aware reinforcement optimization to improve recommendation performance of large language m…
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
Xingjian Wu, Junlin Liu, Xingchen Liu +6
The paper introduces Contrastive Reinforced Policy Optimization (CRPO), a method that frames on‑policy self‑distillation for large language models as a contrastive learning problem…
Operationally Guided Placement-Aware Learning for Industrial Online 3D Bin Packing
Dheeraj Poolavaram, Aanchal Rajesh Chugh, Sebastian Dorn
The paper presents OPAL, a framework that improves industrial online 3D bin packing by using an operationally guided empty-maximal-space generator and a placement-aware learning po…
One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
Rui Tang, Wentao Yang, Peirong Zhang +4
The paper introduces a vision-centric framework called SPaTS that uses a single visual token per text instance and reinforcement learning to improve scene text spotting accuracy an…
MemHarness: Memory Is Reconstructed, Not Replayed
Rong Wu, Daocheng Fu, Licheng Wen +10
The paper introduces MemHarness, a framework that lets large language model agents reconstruct and adapt retrieved past experiences to the current context instead of replaying them…
HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks
Tiangang Li, Xiangbo Tian
The paper introduces HARGO, a reinforcement‑learning post‑training method that weights responses by confidence and reward contrast to better align large language models with divers…
Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold
Songshuo Lu, Zhi Chen, Yaohua Tang
The paper proposes an expand‑then‑compress framework that builds a diverse set of RL‑trained teacher models and then distills them into a single student model, improving reasoning,…
Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
Takumi Shioda, Kohei Terashima, Tatsuo Nagai
The paper investigates using a reasoning language model and TD3-guided reinforcement fine‑tuning to control multi‑zone variable‑air‑volume HVAC systems without building‑specific tr…
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
Haoqing Wang, Xingrun Xing, Wei Xia +2
FaithEyes proposes a multi‑agent framework where a vision‑language model judges its own tool calls to ensure they are useful, improving both accuracy and tool faithfulness on visua…
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
Hongyu Chen, Liang Lin, Guangrun Wang
The paper proposes Self‑Verifying Refinement (SVR), a reinforcement‑learning framework that lets language models decide when to stop refining answers by using their own correctness…
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Junlin Yang, Che Jiang, Yu Fu +21
The paper presents Frontis-MA1, a 35‑billion‑parameter model trained as a meta‑evolution agent for machine learning engineering, using a new OpenMLE stack that combines operator le…
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
Lizhi Yang, Junheng Li, Aaron D. Ames
The paper introduces PAC-MAN, a perception-aware control-barrier-function reinforcement learning framework that enables a humanoid robot to safely dodge balls using depth segmentat…
Harness-G: A Graph-Structured Harness for Search Agents
Yanning Hou, Haoyuan Chen, Sihang Zhou +7
The paper introduces Harness-G, a graph-structured retrieval framework that turns free-form query generation into finite action selection for reinforcement learning search agents a…
RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
Shaobo Liu, Feiqiao Mao, Shuaishuai Zhou +4
RefineSVG introduces a closed-loop visual feedback system that lets large multimodal language models iteratively correct SVG code by rendering the output, comparing it to the targe…
Class-Aware Reinforcement Learning for Counterfactual Explanation Generation
Muhammad Adil Saleem, Syed Ali Raza, Mary-Anne Williams
The paper investigates adding the predicted class of an instance to the reinforcement‑learning state representation for generating counterfactual explanations, showing that this cl…
VIG-RL: Learning to Search and Insert for Verified Image Grounding
Qinhan Yu, Jun Guang, Chong Chen +1
The paper introduces VIG-RL, a reinforcement‑learning based agent that dynamically decides when to retrieve, select, and insert authentic images into text responses, improving veri…
Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Yash Pandya, Sahil Gupta, Sarthak Harne +10
Echoverse introduces a pipeline that compiles specifications into deep, stateful synthetic applications for training computer-use agents, using a co‑evolution loop that repairs env…
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Qixun Wang, Yang Shi, Letian Cheng +11
The paper introduces Beacon, an agentic visual reasoning system that learns when to invoke external tools and how to use them effectively, improving multimodal large language model…
Can Vision-Language Models Reason about AI Edits in Images?
Darsha Udayanga, Pin-Yu Chen, Payel Das +1
The paper explores training vision-language models with reinforcement learning to detect and localize AI-generated image edits, using reasoning traces and a lightweight segmentatio…
LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
Shuang Liang, Haoyang Zhou, Yifan Gong +2
The paper introduces LEEPS, a latent-guided explore‑exploit prompt sampler that selects prompts before rollout to reduce wasted generation budget and improve reinforcement learning…
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Hanzhang Zhou, Panrong Tong, Xu Zhang +13
The paper introduces Qwen-UI-Agent, a foundation model for GUI agents that can operate across mobile, desktop, web, and search environments, combining GUI actions with CLI commands…
X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching
Tianyu Yang, Yiming Zeng, Wenzhe Cai +5
The paper introduces X-NavDP, a diffusion-based visual navigation policy that is fine‑tuned with a novel Group Q-score Reweighted Matching (GQRM) reinforcement learning framework t…
Policy Gradient Steering: Interventions from Behavioral Objectives
Yoann Poupart, Aurélie Beynier, Nicolas Maudet
The paper introduces Policy Gradient Steering (PGS), a reinforcement‑learning based method that uses temporary behavioral objectives to compute removable task vectors for steering…