From the 1 of 6 linked papers with an AI index.
6 papers
Building Agent Harnesses for Scientific Curation from Multimodal Sources
Sheng Zhang, Qin Liu, Renqian Luo +9
The paper introduces Beaver, an agent harness that extracts structured scientific information from papers by integrating text, tables, and figures while preserving provenance, and…
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Qianchu Liu, Sheng Zhang, Guanghui Qin +16
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare appli…
RAP: Runtime Adaptive Pruning for LLM Inference
Huanrong Liu, Chunlin Tian, Xuyang Wei +2
Large language models (LLMs) excel at language understanding and generation, but their enormous computational and memory requirements hinder deployment. Compression offers a potent…
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
James Y. Huang, Wenxuan Zhou, Nan Xu +5
The ability of Large Language Models (LLMs) to generate structured outputs that follow arbitrary schemas is crucial to a wide range of downstream tasks that require diverse structu…
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
Qin Liu, Jacob Dineen, Yuxi Huang +4
Benchmarks are central to measuring the capabilities of large language models and guiding model development, yet widespread data leakage from pretraining corpora undermines their v…
Exploring Scaling Laws for EHR Foundation Models
Sheng Zhang, Qin Liu, Naoto Usuyama +3
The emergence of scaling laws has profoundly shaped the development of large language models (LLMs), enabling predictable performance gains through systematic increases in model si…