9 papers
The Orchestration Gap: Why Process Automation Stalls in Operationally Complex Industries
Jiechao Gao, Yuandong Pan. Yuangang Li, Jie Wang +2
Agentic systems have advanced quickly on digitally native tasks, yet they have barely touched the industries where coordinated automation could matter most: logistics, healthcare o…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
Yuangang Li, Justin Tian Jin Chen, Ethan Yu +2
Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning eva…
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
Jiechao Gao, Rohan Kumar Yadav, Yuangang Li +4
Pretrained language models (PLMs) like BERT provide strong semantic representations but are costly and opaque, while symbolic models such as the Tsetlin Machine (TM) offer transpar…
Mitigating Hallucinations in Large Language Models via Causal Reasoning
Yuangang Li, Yiqing Shen, Yi Nian +7
Large language models (LLMs) exhibit logically inconsistent hallucinations that appear coherent yet violate reasoning principles, with recent research suggesting an inverse relatio…
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
Tiankai Yang, Yi Nian, Shawn Li +9
Anomaly detection (AD) is an important machine learning task with many real-world uses, including fraud detection, medical diagnosis, and industrial monitoring. Within natural lang…