Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Subject-level Inference for Realistic Text Anonymization Evaluation
Myeong Seok Oh, Dong-Yun Kim, Hanseok Oh +6
Current text anonymization evaluation relies on span-based metrics that fail to capture what an adversary could actually infer, and assumes a single data subject, ignoring multi-su…
cs.CL2024
KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
Xiaonan Wang, Jinyoung Yeo, Joon-Ho Lim +1
Large language models have exhibited significant enhancements in performance across various tasks. However, the complexity of their evaluation increases as these models generate mo…