4 papers
Automating Expert-Level Medical Reasoning Evaluation of Large Language Models
Shuang Zhou, Wenya Xie, Jiaxi Li +16
As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluat…
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
An Luo, Xun Xian, Jin Du +12
Large language models (LLMs) have advanced the automation of data science workflows. Yet it remains unclear whether they can critically leverage external domain knowledge as human…
An Outlook on the Opportunities and Challenges of Multi-Agent AI Systems
Fangqiao Tian, An Luo, Jin Du +12
A multi-agent AI system (MAS) is composed of multiple autonomous agents that interact, exchange information, and make decisions based on internal generative models. Recent advances…
EPEE: Towards Efficient and Effective Foundation Models in Biomedicine
Zaifu Zhan, Shuang Zhou, Huixue Zhou +2
Foundation models, including language models, e.g., GPT, and vision models, e.g., CLIP, have significantly advanced numerous biomedical tasks. Despite these advancements, the high…