works on

From the 1 of 14 linked papers with an AI index.

collaborators

14 papers

cs.AI2026

MemoHarness: Agent Harnesses That Learn from Experience

Yue Huang, Wenjie Wang, Han Bao +7

MemoHarness is a framework that automatically adapts the control layer (harness) of large language model agents by learning from past executions, using a dual‑layer experience bank…

cs.AI2026

SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data

Wenjie Wang, Yue Huang, Zhengqing Yuan +6

As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no longer governed by a single universal notion of safety or helpfulness, but ins…

cs.SE2026

UXBench: Measuring the Actionability of LLM-Generated UX Critiques

Wenjie Wang, Yue Huang, Zipeng Ling +11

Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs. Yet no controlled benchmark measures…

cs.AI2026

Confidence Laundering in Agent Systems: Why Uncertainty Needs a Latent Carrier

Kaiwen Shi, Zheyuan Zhang, Han Bao +2

Modern agent systems can turn uncertainty into overconfidence. Fragile upstream decisions are often exposed to downstream components as clean intermediate artifacts, while the unce…

cs.LG2026

Alignment Risks from Capability-Seeking RL Training

Yujun Zhou, Yue Huang, Han Bao +8

While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerabl…

cs.LG2026

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization

Zheyuan Zhang, Kaiwen Shi, Han Bao +3

Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from model-generated outputs but l…