7 papers
Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics
Jiani He, Dingyan Shang, Yihua Xu +4
Reverse-logistics operators often decide how to inspect and route returned assets before their condition is fully observed, while full inspection consumes scarce labor. Semantic Si…
DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains
Shiqi Huang, Jiani He, Dingyan Shang +4
Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value. We present CriticalSCM-Bench v1, a controlled…
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Youting Wang, Xiao Han, Dingyan Shang +2
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentD…
Accuracy-Preserving Stability Regularization for Large-Scale Retail Demand Forecasting
Jize Li, Jiani He, Dishu Yang +3
Retail demand forecasts are reused across replenishment, capacity, labor, and transportation planning cycles. Point-error objectives do not constrain abrupt movement between adjace…
Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors
Bonan Shen, Dingyan Shang, Youting Wang +2
Large language model (LLM) tutors may have access to teacher notes, answer keys, rubrics, or retrieved solutions while producing student-facing explanations. We study whether trunc…
Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking
Bonan Shen, Youting Wang, Dingyan Shang +1
Implicit reward hacking is hard to audit when a language model's chain of thought appears benign: a final answer may be anchored by a prompt shortcut while the written reasoning st…