Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming
Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson +13
We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based test…
cs.CL2025
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
Yining She, Daniel W. Peterson, Marianne Menglin Liu +4
With the increasing adoption of large language models (LLMs), ensuring the safety of LLM systems has become a pressing concern. External LLM-based guardrail models have emerged as…