6 papers
When Not to Write Memory: Governing False Promotion from Correlated Agent Traces
Yijiashun Qi, Xiang Xu, Yuxuan Li
Long-lived language agents increasingly write reusable memories from their own execution traces. The key safety question is not only what agents should remember, but when they shou…
Healthcare Mechanisms from Policy-as-Code Search under Strategic Provider Response
Zihan Wang, Xiang Xu, Hongyuan Zha +1
Healthcare mechanisms are inseparable from the strategic provider response they induce: existing healthcare AI benchmarks hold this response fixed and so cannot evaluate mechanisms…
MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System
Yuliang Xu, Xiang Xu, Yao Wan +2
Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's ability to perform structured r…
When Corrective Hints Hurt: Prompt Design in Reasoner-Guided Repair of LLM Overcaution on Entailed Negations under OWL~2~DL
Yijiashun Qi, Xiang Xu, Yuxuan Li
We report a reproducible error pattern in GPT-5.4 on OWL~2~DL compliance queries: the model frequently answers ``unknown'' when the reasoner-entailed answer is ``no'' under \emph{F…
Beyond Benchmarks: The Economics of AI Inference
Boqin Zhuang, Jiacheng Qiao, Mingqian Liu +8
The inference cost of Large Language Models (LLMs) has become a critical factor in determining their commercial viability and widespread adoption. This paper introduces a quantitat…
WiNGPT-3.0 Technical Report
Boqin Zhuang, Chenxiao Song, Huitong Lu +10
Current Large Language Models (LLMs) exhibit significant limitations, notably in structured, interpretable, and verifiable medical reasoning, alongside practical deployment challen…