1 citations · 1 across the 12 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Accelerating Speculative Decoding with Block Diffusion Draft Trees
Liran Ringel, Yaniv Romano
Speculative decoding accelerates autoregressive language models by using a lightweight drafter to propose multiple future tokens, which the target model then verifies in parallel.…
cs.CL2025
Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
Liran Ringel, Elad Tolochinsky, Yaniv Romano
Test-time scaling has emerged as an effective approach for improving language model performance by utilizing additional compute at inference time. Recent studies have shown that ov…
cs.CL2024★ 1 cited
Segment-Based Attention Masking for GPTs
Shahar Katz, Liran Ringel, Yaniv Romano +1
Modern Language Models (LMs) owe much of their success to masked causal attention, the backbone of Generative Pre-Trained Transformer (GPT) models. Although GPTs can process the en…