collaborators

16 papers

cs.CL2026

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

Ajmal M., Abin Roy, Afthab Salam Kanniyan +4

Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation…

cs.AI2026

Rethinking Uncertainty Evaluation in Large Language Models

Krish Matta, Atharv Naphade, Andy Zou

Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and d…

cs.CY2026

On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

Yue Huang, Chujie Gao, Siyuan Wu +63

Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions.…

cs.CY2026

Muse Spark Safety & Preparedness Report

Cristina Menghini, Peter Ney, Hamza Kwisaba +117

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…

cs.CL2026

TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models

Jinho Choo, JunSeung Lee, Jimyeong Kim +3

Large language models (LLMs) demonstrate strong multilingual capabilities, yet often fail to consistently generate responses in the intended language, exhibiting a phenomenon known…

cs.CL2026

On Safety Risks in Experience-Driven Self-Evolving Agents

Weixiang Zhao, Yichen Zhang, Yingshuo Wang +8

Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduc…