collaborators

5 papers

cs.AI2026

MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models

Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian +1

Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms…

cs.AI2026

Designing for Doubt: The Case for Informed Abstention in Autonomous Agents

Victor Ojewale, Suresh Venkatasubramanian

As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate…

cs.CL2026

Multi-lingual Functional Evaluation for Large Language Models

Victor Ojewale, Inioluwa Deborah Raji, Suresh Venkatasubramanian

Multi-lingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. However, these evaluations often fail to provide…

cs.CY2026

Audit Trails for Accountability in Large Language Models

Victor Ojewale, Harini Suresh, Suresh Venkatasubramanian

Large language models (LLMs) are increasingly embedded in consequential decisions across healthcare, finance, employment, and public services. Yet accountability remains fragile be…

cs.CY2025

Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling

Victor Ojewale, Ryan Steed, Briana Vecchione +2

Audits are critical mechanisms for identifying the risks and limitations of deployed artificial intelligence (AI) systems. However, the effective execution of AI audits remains inc…