activity
20242026
collaborators

7 papers

cs.LG2026

Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges

Chen Feng, Minghe Shen, Ananth Balashankar +2

Reliable certification of Large Language Models (LLMs)-verifying that failure rates are below a safety threshold-is critical yet challenging. While "LLM-as-a-Judge" offers scalabil…

cs.CL2025

Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications

Xiao Ye, Jacob Dineen, Zhaonan Li +11

Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a…

cs.SE2025

TALM: Dynamic Tree-Structured Multi-Agent Framework with Long-Term Memory for Scalable Code Generation

Ming-Tung Shen, Yuh-Jzer Joung

Agentic code generation requires large language models (LLMs) capable of complex context management and multi-step reasoning. Prior multi-agent frameworks attempt to address these…

cs.CL2025

CC-LEARN: Cohort-based Consistency Learning

Xiao Ye, Shaswat Shrivastava, Zhaonan Li +6

Large language models excel at many tasks but still struggle with consistent, robust reasoning. We introduce Cohort-based Consistency Learning (CC-Learn), a reinforcement learning…

cs.CL2025

QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA

Jacob Dineen, Aswin RRV, Qin Liu +8

Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the tra…

cs.AI2025

Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development

Ming Shen, Raphael Shu, Anurag Pratik +4

We have seen remarkable progress in large language models (LLMs) empowered multi-agent systems solving complex tasks necessitating cooperation among experts with diverse skills. Ho…