collaborators

10 papers

cs.AI2026

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes…

cs.SE2026

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation

Jie "JW" Wu

Large language models (LLMs) generate code from natural-language prompts, yet real-world prompts rarely provide complete specifications. When prompts leave input formats, error han…

cs.SE2026

GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents

Jie JW Wu, Ayanda Patrick Herlihy, Ahmad Saleem Mirza +2

With data-driven development now widely adopted, online A/B testing is an established method for measuring the effects of new technologies. However, deploying online experiments de…

cs.AI2026

AnyEdit++: Adaptive Long-Form Knowledge Editing via Bayesian Surprise

Bowen Tian, Caixue He, Jiemin Wu +4

Editing complex, long-form knowledge in Large Language Models remains a significant challenge due to the difficulty of maintaining generation coherence. Existing autoregressive met…

cs.CL2026

Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning

Xuewei Yang, Jiachen Yu, Jie Wu +3

Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in which increasingly concentrated…

cs.SE2025

MANTRA: a Framework for Multi-stage Adaptive Noise TReAtment During Training

Zixiao Zhao, Fatemeh H. Fard, Jie JW Wu

The reliable application of deep learning models to software engineering tasks hinges on high-quality training data. Yet, large-scale repositories inevitably introduce noisy or mis…