4 papers
Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense
Minseok Choi, Seungbin Yang, Dongjin Kim +5
Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evol…
K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance
Eunbyeol Cho, Yunseung Lee, Mirae Kim +3
Large Language Models (LLMs) have advanced financial automation through Retrieval-Augmented Generation (RAG), yet hallucinations remain a critical barrier to deployment in high-sta…
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
Yunseung Lee, Subin Kim, Youngjun Kwak +1
Large language models (LLMs)-based chatbots are increasingly being adopted in the financial domain, particularly in digital banking, to handle customer inquiries about products suc…
FinNuE: Exposing the Risks of Using BERTScore for Numerical Semantic Evaluation in Finance
Yu-Shiang Huang, Yun-Yu Lee, Tzu-Hsin Chou +2
BERTScore has become a widely adopted metric for evaluating semantic similarity between natural language sentences. However, we identify a critical limitation: BERTScore exhibits l…