Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, cr…
cs.AI2026
The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure
Yubo Li, Ramayya Krishnan, Rema Padman
Reasoning models are evaluated on single-turn benchmarks but deployed in multi-turn dialogue, where users push back on correct answers. Under sustained adversarial pressure we find…
cs.AI2024
Reliability, Resilience and Human Factors Engineering for Trustworthy AI Systems
Saurabh Mishra, Anand Rao, Ramayya Krishnan +3
As AI systems become integral to critical operations across industries and services, ensuring their reliability and safety is essential. We offer a framework that integrates establ…