works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.AI2026

Building Agent Harnesses for Scientific Curation from Multimodal Sources

Sheng Zhang, Qin Liu, Renqian Luo +9

The paper introduces Beaver, an agent harness that extracts structured scientific information from papers by integrating text, tables, and figures while preserving provenance, and…

cs.AI2026

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

Qianchu Liu, Sheng Zhang, Guanghui Qin +16

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare appli…

cs.LG2026

RAP: Runtime Adaptive Pruning for LLM Inference

Huanrong Liu, Chunlin Tian, Xuyang Wei +2

Large language models (LLMs) excel at language understanding and generation, but their enormous computational and memory requirements hinder deployment. Compression offers a potent…

cs.CL2025

OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas

James Y. Huang, Wenxuan Zhou, Nan Xu +5

The ability of Large Language Models (LLMs) to generate structured outputs that follow arbitrary schemas is crucial to a wide range of downstream tasks that require diverse structu…

cs.CL2025

ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation

Qin Liu, Jacob Dineen, Yuxi Huang +4

Benchmarks are central to measuring the capabilities of large language models and guiding model development, yet widespread data leakage from pretraining corpora undermines their v…

cs.CL2025

Exploring Scaling Laws for EHR Foundation Models

Sheng Zhang, Qin Liu, Naoto Usuyama +3

The emergence of scaling laws has profoundly shaped the development of large language models (LLMs), enabling predictable performance gains through systematic increases in model si…