3 papers
cs.AI2026
Streaming Hallucination Detection in Long Chain-of-Thought Reasoning
Haolang Lu, Minghui Pan, Ripeng Li +6
Long chain-of-thought (CoT) reasoning improves the performance of large language models, yet hallucinations in such settings often emerge subtly and propagate across reasoning step…
cs.CR2025
LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
Yuanhe Zhang, Weiliu Wang, Zhenhong Zhou +5
Large Language Model (LLM)-based agents have demonstrated remarkable capabilities in reasoning, planning, and tool usage. The recently proposed Model Context Protocol (MCP) has eme…
cs.CL2025
LIFEBench: Evaluating Length Instruction Following in Large Language Models
Wei Zhang, Zhenhong Zhou, Kun Wang +9
While large language models (LLMs) can solve PhD-level reasoning problems over long context inputs, they still struggle with a seemingly simpler task: following explicit length ins…