3 papers
cs.AI2026
ForecastBench-Sim: A Simulated-World Forecasting Benchmark
Jaeho Lee, Nick Merrill, Ezra Karger
Forecasting benchmarks for general-purpose AI systems usually inherit the constraints of the real world: outcomes resolve slowly, tail events are rare, and counterfactual questions…
cs.CR2026
Symmetry Defeats Auditing
Nick Merrill, Zeke Medley
We demonstrate an attack on Introspection Adapters (Shenoy et al., 2026).
cs.AI2026
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
Nick Merrill, Jaeho Lee, Ezra Karger
We document inverse scaling in LLMs on forecasting problems whose underlying time series exhibit superlinear growth and tail risk of regime change, a structure common in finance an…