collaborators

5 papers

cs.AI2026

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision

Wei-Jung Huang, Bonan Shen

LLM-agent evaluations often produce task outcomes long before the full benchmark run is complete. A partial score is tempting to report, but it does not show whether the observed t…

cs.LG2026

AdaStop: Cost-Aware Early Stopping for DNN Test Selection

Bonan Shen, Wei-Jung Huang, Xin Liu +2

Existing methods for testing deep neural networks (DNNs) primarily prioritize test inputs likely to reveal model faults under a fixed labeling budget. In practice, choosing that bu…

cs.SE2026

LLM-Driven CI-CD Workflow Intelligence for Cyber Systems Engineering

Bonan Shen, Jiazhou Gao, Tao Ning +2

CI/CD workflows have become executable operational policy: they decide what gets built, tested, released, and deployed, and they mediate how maintainers interact with delivery infr…

cs.AI2026

Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors

Bonan Shen, Dingyan Shang, Youting Wang +2

Large language model (LLM) tutors may have access to teacher notes, answer keys, rubrics, or retrieved solutions while producing student-facing explanations. We study whether trunc…

cs.AI2026

Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking

Bonan Shen, Youting Wang, Dingyan Shang +1

Implicit reward hacking is hard to audit when a language model's chain of thought appears benign: a final answer may be anchored by a prompt shortcut while the written reasoning st…