autonomous multi-agent systems 1black-box extraction 1defense mechanisms 1inference-time harness 1ip leakage 1llm security 1
From the 1 of 13 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
Hossein Entezari Zarch, Lei Gao, Chaoyi Jiang +1
Large reasoning models (LRMs) achieve state-of-the-art performance on challenging benchmarks by generating long chains of intermediate steps, but their inference cost is dominated…
cs.CL2025
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
Hossein Entezari Zarch, Lei Gao, Chaoyi Jiang +1
Speculative Decoding (SD) is a widely used approach to accelerate the inference of large language models (LLMs) without reducing generation quality. It operates by first using a co…