Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Compared to What? Baselines and Metrics for Counterfactual Prompting
Zihao Yang, Mosh Levy, Yoav Goldberg +1
Counterfactual prompting (i.e., perturbing a single factor and measuring output change) is widely used to evaluate things like LLM bias and CoT faithfulness. But in this work we ar…
cs.CL2025
State over Tokens: Characterizing the Role of Reasoning Tokens
Mosh Levy, Zohar Elyoseph, Shauli Ravfogel +1
Large Language Models (LLMs) can generate reasoning tokens before their final answer to boost performance on complex tasks. While these sequences seem like human thought processes,…
cs.CL2024
Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
Mosh Levy, Alon Jacoby, Yoav Goldberg
This paper explores the impact of extending input lengths on the capabilities of Large Language Models (LLMs). Despite LLMs advancements in recent times, their performance consiste…