13 papers
Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens
Peizhi Niu, Wenjie Qu, Shangding Gu +14
Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on system-level responsibilities…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
LLMs Should Express Uncertainty Explicitly
Junyu Guo, Shangding Gu, Ming Jin +2
Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We study whether post-training can make a m…
StyleBench: Evaluating thinking styles in Large Language Models
Junyu Guo, Shangding Gu, Ming Jin +2
Structured reasoning can improve the inference performance of large language models (LLMs), but it also introduces computational cost and control constraints. When additional reaso…
Representation Learning Enhanced Deep Reinforcement Learning for Optimal Operation of Hydrogen-based Multi-Energy Systems
Zhenyu Pu, Yu Yang, Lun Yang +3
Hydrogen-based multi-energy systems (HMES) have emerged as a promising low-carbon and energy-efficient solution, as it can enable the coordinated operation of electricity, heating…
An Equivalent and Unified Virtual Battery Modeling Framework for Flexibility Characterization of Building HVAC Systems
Qi Zhu, Yu Yang, Liang Yu +3
The heating, ventilation and air-conditioning (HVAC) system dominates building's energy consumption and meanwhile exhibits substantial operational flexibility that can be exploited…