2 papers
cs.CR2026
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
Zhicheng Fang, Jingjie Zheng, Chenxu Fu +1
Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datase…
cs.CL2026
QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models
Zhaolu Kang, Junhao Gong, Wenqing Hu +15
Large Language Models (LLMs) have shown strong capabilities across many domains, yet their evaluation in financial quantitative tasks remains fragmented and mostly limited to knowl…