2 papers
cs.LG2025
Enhancing PPO with Trajectory-Aware Hybrid Policies
Qisai Liu, Zhanhong Jiang, Hsin-Jung Yang +3
Proximal policy optimization (PPO) is one of the most popular state-of-the-art on-policy algorithms that has become a standard baseline in modern reinforcement learning with applic…
cs.CV2025
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception
Joshua R. Waite, Md. Zahid Hasan, Qisai Liu +3
Vision-language model (VLM) fine-tuning for application-specific visual grounding based on natural language instructions has become one of the most popular approaches for learning-…