4 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.AI2025★ 1 cited
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
Clinton J. Wang, Dean Lee, Cristina Menghini +7
As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging mu…
cs.LG2024★ 4 cited
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
Nathaniel Li, Ziwen Han, Ian Steneker +6
Recent large language model (LLM) defenses have greatly improved models' ability to refuse harmful queries, even when adversarially attacked. However, LLM defenses are primarily ev…