5 papers
Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
Jiayu Wang, Weijiang Lv, Bowen Fu +8
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and eve…
Scientific judgment drifts over time in AI ideation
Lingyu Zhang, Mitchell Wang, Boyuan Chen
Scientific discovery begins with ideas, yet evaluating early-stage research concepts is a subtle and subjective human judgment. As large language models (LLMs) are increasingly tas…
Enabling Multi-Robot Collaboration from Single-Human Guidance
Zhengran Ji, Lingyu Zhang, Paul Sajda +1
Learning collaborative behaviors is essential for multi-agent systems. Traditionally, multi-agent reinforcement learning solves this implicitly through a joint reward and centraliz…
CREW: Facilitating Human-AI Teaming Research
Lingyu Zhang, Zhengran Ji, Boyuan Chen
With the increasing deployment of artificial intelligence (AI) technologies, the potential of humans working with AI agents has been growing at a great speed. Human-AI teaming is a…
GUIDE: Real-Time Human-Shaped Agents
Lingyu Zhang, Zhengran Ji, Nicholas R Waytowich +1
The recent rapid advancement of machine learning has been driven by increasingly powerful models with the growing availability of training data and computational resources. However…