Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
Tianyu Liu, Allen Xin Wang, Antonia Panescu +30
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchma…
cs.AI2024
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
Xinyang Hu, Fengzhuo Zhang, Siyu Chen +1
Chain-of-Thought (CoT) prompting and its variants have gained popularity as effective methods for solving multi-step reasoning problems using pretrained large language models (LLMs…