papers

Publications (20)

cs.LG2025

SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph

Jiazheng Li, Yawei Wang, David Yan +5

Large Language Models (LLMs) have demonstrated remarkable capabilities, enabling language agents to excel at single-turn tasks. However, their application to complex, multi-step, a…

cs.LG2021

Reward function shape exploration in adversarial imitation learning: an empirical study

Yawei Wang, Xiu Li

For adversarial imitation learning algorithms (AILs), no true rewards are obtained from the environment for learning the strategy. However, the pseudo rewards based on the output o…

cs.LG2020

Wasserstein Distance guided Adversarial Imitation Learning with Reward Shape Exploration

Ming Zhang, Yawei Wang, Xiaoteng Ma +4

The generative adversarial imitation learning (GAIL) has provided an adversarial learning framework for imitating expert policy from demonstrations in high-dimensional continuous t…

cs.CL2026

RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation

Zhichao Xu, Minheng Wang, Yawei Wang +4

Search agents trained with reinforcement learning (RL) interleave reasoning with tool calls in a multi-turn, tool-integrated reasoning (TIR) loop, where each tool invocation return…

physics.app-ph2019

Fully superconducting machine for electric aircraft propulsion: study of AC loss for HTS stator

Fangjing Weng, Min Zhang, Tian Lan +2

Fully superconducting machines provide the high power density required for future electric aircraft propulsion. However, superconducting windings generate AC losses in AC electrica…

cs.LG2025

Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs

Dongkyu Derek Cho, Huan Song, Arijit Ghosh Chowdhury +6

Fine-tuning large language models (LLMs) for downstream tasks typically exhibit a fundamental safety-capability tradeoff, where improving task performance degrades safety alignment…

cs.AI2026

A Sober Look at Agentic Misalignment in Automated Workflows

Wenqian Ye, Bo Yuan, Zhichao Xu +4

We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to agentic misalignment. Although these systems can solv…

cs.AI2026

Reinforcement Learning for Self-Improving Agent with Skill Library

Jiongxiao Wang, Qiaojing Yan, Yawei Wang +6

Large Language Model (LLM)-based agents have demonstrated remarkable capabilities in complex reasoning and multi-turn interactions but struggle to continuously improve and adapt wh…

cs.LG2025

Uncovering Causal Relation Shifts in Event Sequences under Out-of-Domain Interventions

Kazi Tasnim Zinat, Yun Zhou, Xiang Lyu +3

Inferring causal relationships between event pairs in a temporal sequence is applicable in many domains such as healthcare, manufacturing, and transportation. Most existing work on…

cs.AI2026

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning

Wenhao Zhang, Yibo Xie, Rui Wang +9

Autoregressive rollout generation is a major computational cost in reinforcement learning for large language models. Reusing each rollout batch for additional learner updates amort…

cs.CV2025

SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces

Guande Wu, Huan Song, Yawei Wang +4

Reasoning is increasingly crucial for various tasks. While chain-of-thought prompting enables large language models to leverage reasoning effectively, harnessing the reasoning capa…

cond-mat.supr-con2019

3D quench modeling based on T-A formulation for high temperature superconductor CORC cables

Yawei Wang, Jinxing Zheng, Zixuan Zhu +2

High temperature superconductor (HTS) (RE)Ba2Cu3Ox (REBCO) conductor on round core cable (CORC) has high current carrying capacity for high field magnet and power applications. In…

cs.CR2024

A Systematic Survey of Blockchained Federated Learning

Zhilin Wang, Qin Hu, Minghui Xu +3

With the technological advances in machine learning, effective ways are available to process the huge amount of data generated in real life. However, issues of privacy and scalabil…

physics.bio-ph2024

Hybrid roles of adaptation and optimization in formation of vascular network

Yawei Wang, Zilu Qin, Yubo Fan

It was hypothesized that the structures of biological transport networks are the result of either energy consumption or adaptation dynamics. Although approaches based on these hypo…

cs.AI2026

CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection

Linbo Liu, Guande Wu, Han Ding +7

Large language model agents rely on effective model context to obtain task-relevant information for decision-making. Many existing context engineering approaches primarily rely on…

eess.SY2020

Autonomous Charging of Electric Vehicle Fleets to Enhance Renewable Generation Dispatchability

Reza Bayani, Saeed D. Manshadi, Guangyi Liu +2

A total 19% of generation capacity in California is offered by PV units and over some months, more than 10% of this energy is curtailed. In this research, a novel approach to reduc…

cs.SE2021

Diversity-aware Web APIs Recommendation with Compatibility Guarantee

Wenwen Gonga, Yulan Zhang, Xuyun Zhang +4

With the ever-increasing prevalence of web APIs (Application Programming Interfaces) in enabling smart software developments, finding and composing a list of existing web APIs that…

cs.CL2025

A Systematic Survey of Automatic Prompt Optimization Techniques

Kiran Ramnath, Kang Zhou, Sheng Guan +18

Since the advent of large language models (LLMs), prompt engineering has been a crucial step for eliciting desired responses for various Natural Language Processing (NLP) tasks. Ho…

cs.SI2021

Diversified and Compatible Web APIs Recommendation in IoT

Wenwen Gong, Huiping Wu, Xiaokang Wang +4

With the ever-increasing popularity of Service-oriented Architecture (SoA) and Internet of Things (IoT), a considerable number of enterprises or organizations are attempting to enc…

cs.CL2025

Reinforcement Mid-Training

Yijun Tian, Shaoyu Chen, Zhichao Xu +4

The development of state-of-the-art large language models is commonly understood as a two-stage process involving pre-training and post-training. We point out the need for an addit…