activity
20242026
collaborators

7 papers

cs.CR2026

Towards Compositional Generalization in LLMs for Smart Contract Security: A Case Study on Reentrancy Vulnerabilities

Ying Zhou, Jiacheng Wei, Yu Qi +2

Large language models (LLMs) demonstrate remarkable capabilities in natural language understanding and generation. Despite being trained on large-scale, high-quality data, LLMs sti…

cs.LG2025

Preference-Guided Reinforcement Learning for Efficient Exploration

Guojian Wang, Jianxiang Liu, Xinyuan Li +4

In this paper, we investigate preference-based reinforcement learning (PbRL), which enables reinforcement learning (RL) agents to learn from human feedback. This is particularly va…

cs.CL2025

A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs

Jiacheng Wei, Faguo Wu, Xiao Zhang

Large language models are unable to continuously adapt and learn from new data during reasoning at inference time. To address this limitation, we propose that complex reasoning tas…

cs.LG2025

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood

Qingmao Yao, Zhichao Lei, Tianyuan Chen +5

Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the -value overestimation for out-of-distribution (OOD) actions. Existing methods address th…

cs.LG2024

Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations

Guojian Wang, Faguo Wu, Xiao Zhang +1

The sparsity of reward feedback remains a challenging problem in online deep reinforcement learning (DRL). Previous approaches have utilized offline demonstrations to achieve impre…

cs.LG2024

Trajectory-Oriented Policy Optimization with Sparse Rewards

Guojian Wang, Faguo Wu, Xiao Zhang

Mastering deep reinforcement learning (DRL) proves challenging in tasks featuring scant rewards. These limited rewards merely signify whether the task is partially or entirely acco…