4 papers
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
Lukas Weidener, Marko BrkiÄ, Mihailo JovanoviÄ +2
Frontier large language models are increasingly deployed as orchestration backbones for biological research workflows, yet no shared evidence base exists for comparing their refusa…
From Agent-Only Social Networks to Autonomous Scientific Research: Lessons from OpenClaw and Moltbook, and the Architecture of ClawdLab and Beach.Science
Lukas Weidener, Marko BrkiÄ, Phillip Lee +3
In January 2026, the open-source agent framework OpenClaw and the agent-only social network Moltbook produced a large-scale dataset of autonomous AI-to-AI interaction, attracting s…
Rethinking the AI Scientist: Interactive Multi-Agent Workflows for Scientific Discovery
Lukas Weidener, Marko BrkiÄ, Mihailo JovanoviÄ +5
Artificial intelligence systems for scientific discovery have demonstrated remarkable potential, yet existing approaches remain largely proprietary and operate in batch-processing…
From Task Executors to Research Partners: Evaluating AI Co-Pilots Through Workflow Integration in Biomedical Research
Lukas Weidener, Marko BrkiÄ, Chiara Bacci +7
Artificial intelligence systems are increasingly deployed in biomedical research. However, current evaluation frameworks may inadequately assess their effectiveness as research col…