Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling
Shubh Chapra, Dhruv Kumar, Murari Mandal +1
We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows. CCB fixes the…
cs.AI2026
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
Devanshu Sahoo, Manish Prasad, Vasudev Majhi +5
Driven by surging submission volumes, scientific peer review has catalyzed two parallel trends: individual over-reliance on LLMs and institutional AI-powered assessment systems. Th…
cs.AI2025
OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs
Debdeep Sanyal, Umakanta Maharana, Yash Sinha +4
Role-based access control (RBAC) and hierarchical structures are foundational to how information flows and decisions are made within virtually all organizations. As the potential o…