works on

From the 1 of 44 linked papers with an AI index.

activity
20242026
collaborators

44 papers

cs.AI2026

AIMO Interpretability Challenge

Michal Štefánik, Philipp Mondorf, Andreas Waldis +11

The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…

cs.AI2026

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

Maheep Chaudhary, Fazl Barez

White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure safe model behavior. However, white-box…

cs.CY2026

The 2026 Singapore Consensus on Global AI Safety Research Priorities

Stephen Casper, Oskar Galeev, Yoshua Bengio +117

Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 202…

cs.LG2026

Pretraining Curricula Enable Selective Fine-tuning

Sebastian A. Bruijns, Jirko Rubruck, Mia H. Whitefield +3

Transformers follow implicit curricula whereby some tasks are learned before others. However, how explicit pretraining curricula influence learning, generalization, and the selecti…

cs.AI2026

The Capability Frontier: Benchmarks Miss 82% of Model Performance

Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8

Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data…

cs.LG2026

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +22

As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchma…