3 papers
cs.AI2026
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
Yibing Liu, Yangze Liu, Xiaolong Yin +4
Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambiguity, unsafe external write…
cs.CL2026
FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents
Haoxuan Jia, Yang Liu, Bin Chong +10
Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary filters often miss irrevers…
cs.SI2026
Reducing Detail Hallucinations in Long-Context Regulatory Understanding via Targeted Preference Optimization
Yang Liu, Bin Chong, Yuhan Lin +7
Large language models (LLMs) frequently produce \emph{detail hallucinations} when processing long regulatory documents, including subtle errors in threshold values, units, scopes,…