5 papers
Forecasting Future Behavior as a Learning Task
Mosh Levy, Yoav Goldberg, Asa Cooper Stickland
Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior on new inputs. For large reasoning models (LRMs), this convent…
Compared to What? Baselines and Metrics for Counterfactual Prompting
Zihao Yang, Mosh Levy, Yoav Goldberg +1
Counterfactual prompting (i.e., perturbing a single factor and measuring output change) is widely used to evaluate things like LLM bias and CoT faithfulness. But in this work we ar…
An Independent Safety Evaluation of Kimi K2.5
Zheng-Xin Yong, Parv Mahajan, Andy Wang +12
Kimi K2.5 is an open-weight LLM that rivals closed models across coding, multimodal, and agentic benchmarks, but was released without an accompanying safety evaluation. In this wor…
State over Tokens: Characterizing the Role of Reasoning Tokens
Mosh Levy, Zohar Elyoseph, Shauli Ravfogel +1
Large Language Models (LLMs) can generate reasoning tokens before their final answer to boost performance on complex tasks. While these sequences seem like human thought processes,…
Humans Perceive Wrong Narratives from AI Reasoning Texts
Mosh Levy, Zohar Elyoseph, Yoav Goldberg
A new generation of AI models generates step-by-step reasoning text before producing an answer. This text appears to offer a human-readable window into their computation process, a…