collaborators

7 papers

cs.AI2026

READY or Not: Reliable Enterprise Agent Deployment

Veronica Chatrath, Bryan Zhu, Jingxuan Fan +15

An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, w…

cs.AI2026

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Veronica Chatrath, Bryan Zhu, George Pu +16

Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, lo…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.LG2026

Scaling Reward Modeling without Human Supervision

Jingxuan Fan, Yueying Li, Zhenting Qi +4

Learning from feedback is an instrumental process for advancing the capabilities and safety of frontier models, yet its effectiveness is often constrained by cost and scalability.…

cs.AI2026

Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformers

Binxu Wang, Jingxuan Fan, Xu Pan

Diffusion Transformers (DiTs) have greatly advanced text-to-image generation, but models still struggle to generate the correct spatial relations between objects as specified in th…

cs.CL2025

Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs

Xu Pan, Ely Hahami, Jingxuan Fan +2

Large language models (LLMs) are often used in environments where facts evolve, yet factual knowledge updates via fine-tuning on unstructured text often suffer from 1) reliance on…