collaborators

8 papers

cs.AI2026

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Yijun Zhang, Fan Xu, Jiaxin Ding +6

Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of interme…

cs.CV2026

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models

Huawei Ji, Yuanhao Sun, Yuan Jin +4

Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in vision-language models (VLMs).…

cs.AI2026

Inductive Reasoning for Temporal Knowledge Graphs with Emerging Entities

Ze Zhao, Yuhui He, Lyuwen Wu +6

Reasoning on Temporal Knowledge Graphs (TKGs) is essential for predicting future events and time-aware facts. While existing methods are effective at capturing relational dynamics,…

cs.CL2026

RADAR: Reasoning as Discrimination with Aligned Representations for LLM-based Knowledge Graph Reasoning

Bo Xue, Yuan Jin, Luoyi Fu +2

Knowledge graph reasoning (KGR) infers missing facts, with recent advances increasingly harnessing the semantic priors and reasoning abilities of Large Language Models (LLMs). Howe…

cs.CL2026

<SOG_k>: One LLM Token for Explicit Graph Structural Understanding

Jingyao Wu, Bin Lu, Zijun Di +5

Large language models show great potential in unstructured data understanding, but still face significant challenges with graphs due to their structural hallucination. Existing app…

cs.LG2026

Extreme Value Policy Optimization for Safe Reinforcement Learning

Shiqing Gao, Yihang Zhou, Shuai Shao +5

Ensuring safety is a critical challenge in applying Reinforcement Learning (RL) to real-world scenarios. Constrained Reinforcement Learning (CRL) addresses this by maximizing retur…