3 papers
cs.CR2026
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
Taein Lim, Seongyong Ju, Munhyeok Kim +2
Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenomenon: different agents exhibit…
cs.SE2025
MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
Hyunjun Kim, Sejong Kim
We introduce MacroBench, a code-first benchmark that evaluates whether LLMs can synthesize reusable browser-automation programs (macros) from natural-language goals by reading HTML…
cs.IR2025
Optimizing Retrieval Strategies for Financial Question Answering Documents in Retrieval-Augmented Generation Systems
Sejong Kim, Hyunseo Song, Hyunwoo Seo +1
Retrieval-Augmented Generation (RAG) has emerged as a promising framework to mitigate hallucinations in Large Language Models (LLMs), yet its overall performance is dependent on th…