15 papers
Open Technical Problems in Open-Weight AI Model Risk Management
Stephen Casper, Kyle O'Brien, Shayne Longpre +19
Frontier AI models with openly available weights are steadily becoming more powerful and widely adopted. However, compared to proprietary models, open-weight models pose different…
Voluntary Collusion with Secret Tools in Competing LLM Agents
Xijie Zeng, Frank Rudzicz
Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret collusion whenever doing so confer…
Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains
Raeid Saqur, Christoph Bergmeir, Blanka Horvath +3
We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong, persistent periodicities and seasonaliti…
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
Huaizhi Ge, Frank Rudzicz, Zining Zhu
Large language models (LLMs) have demonstrated remarkable capabilities, but updating their knowledge post-training remains a critical challenge. While recent model editing techniqu…
Understanding Language Model Circuits through Knowledge Editing
Huaizhi Ge, Frank Rudzicz, Zining Zhu
Recent advances in language model interpretability have identified circuits, critical subnetworks that replicate model behaviors, yet how knowledge is structured within these cruci…
LLM Library Learning Fails: A LEGO-Prover Case Study
Ian Berlot-Attwell, Frank Rudzicz, Xujie Si
Recent advancements in the coding, reasoning, and tool-using abilities of LLMs have spurred interest in library learning (i.e., online learning through the creation, storage, and r…