activity
20242026
collaborators

7 papers

cs.LG2026

Representation over Routing: Diagnosing Temporal Routing Pathologies in Multi-Timescale PPO

Jing Sun

Temporal credit assignment in reinforcement learning is often approached by introducing value estimates at multiple discount factors. A natural next step is to let the actor dynami…

cs.AI2026

AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning

Jingbo Sun, Wenyue Chong, Songjun Tu +7

Agentic retrieval-augmented generation (RAG) systems enable large language models (LLMs) to solve complex tasks through multi-step interaction with external retrieval tools. Howeve…

cs.CV2026

Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning

Jingbo Sun, Qichao Zhang, Songjun Tu +5

Zero-shot unsupervised reinforcement learning (URL) offers a promising direction for building generalist agents capable of generalizing to unseen tasks without additional supervisi…

cs.MM2025

Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization

Songjun Tu, Qichao Zhang, Jingbo Sun +6

While multimodal large language models excel at tasks that integrate visual perception with symbolic reasoning, their performance is often undermined by a critical vulnerability: p…

cs.AI2025

Salience-Invariant Consistent Policy Learning for Generalization in Visual Reinforcement Learning

Jingbo Sun, Songjun Tu, Qichao Zhang +2

Generalizing policies to unseen scenarios remains a critical challenge in visual reinforcement learning, where agents often overfit to the specific visual observations of the train…

cs.LG2024

Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model

Songjun Tu, Jingbo Sun, Qichao Zhang +2

Preference-based reinforcement learning (PbRL) provides a powerful paradigm to avoid meticulous reward engineering by learning rewards based on human preferences. However, real-tim…