Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models
Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian +1
Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms…
cs.AI2026
Designing for Doubt: The Case for Informed Abstention in Autonomous Agents
Victor Ojewale, Suresh Venkatasubramanian
As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate…