12 papers
LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning
Yubo Li, Ramayya Krishnan, Rema Padman
When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defa…
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Ajay Patel, Kartik Hosanagar, Ramayya Krishnan +3
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answe…
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, cr…
Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks
Yubo Li, Rema Padman, Ramayya Krishnan
Health systems are rapidly deploying generative AI assistants that answer patient questions from institution-authored education materials, on the premise that grounding in local co…
The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
Yubo Li, Lu Zhang, Tianchong Jiang +2
Large language models fail when a salient surface cue conflicts with an unstated feasibility constraint. We introduce the Heuristic Override Benchmark (HOB): 500 instances spanning…
Toward Agentic Governance: What Shapes LLM-Agent Intervention in Public Forums?
Luyang Zhang, Yi-Yun Chu, Ramayya Krishnan
LLM agents are increasingly used in moderation-relevant public forum workflows, where their choices to answer, acknowledge, repair, or decline are routinely challenged by users, pl…