3 papers
cs.CR2026
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
Xuyang Liu, Yibin Han, Zhenwei Zhang +8
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions…
cs.DB2026
MMTS-BENCH: A Comprehensive Benchmark for Time Series Understanding and Reasoning
Yao Yin, Zhenyu Xiao, Musheng Li +7
Time series data are central to domains such as finance, healthcare, and cloud computing, yet existing benchmarks for evaluating various large language models (LLMs) on temporal ta…
cs.LG2024
See it, Think it, Sorted: Large Multimodal Models are Few-shot Time Series Anomaly Analyzers
Jiaxin Zhuang, Leon Yan, Zhenwei Zhang +3
Time series anomaly detection (TSAD) is becoming increasingly vital due to the rapid growth of time series data across various sectors. Anomalies in web service data, for example,…