2 papers
cs.CL2026
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
Jiacheng Guo, Suozhi Huang, Zixin Yao +16
This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large Language Model (LLM) agents in t…
cs.AI2025
CRISPR-GPT for Agentic Automation of Gene-editing Experiments
Yuanhao Qu, Kaixuan Huang, Ming Yin +11
The introduction of genome engineering technology has transformed biomedical research, making it possible to make precise changes to genetic information. However, creating an effic…