3 papers
cs.AI2026
ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
Sanghyu Yoon, Dongmin Kim, Suhee Yoon +6
In tabular anomaly detection (AD), textual semantics often carry critical signals, as the definition of an anomaly is closely tied to domain-specific context. However, existing ben…
cs.CL2026
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
Seungdong Yoa, Sanghyu Yoon, Suhee Yoon +4
The evaluation of large language models (LLMs) has predominantly relied on static datasets, which offer limited scalability and fail to capture the evolving reasoning capabilities…
cs.LG2025
MultiTab: A Comprehensive Benchmark Suite for Multi-Dimensional Evaluation in Tabular Domains
Kyungeun Lee, Moonjung Eo, Hye-Seung Cho +5
Despite the widespread use of tabular data in real-world applications, most benchmarks rely on average-case metrics, which fail to reveal how model behavior varies across diverse d…