2 papers
cs.CL2026
Message Passing Enables Efficient Reasoning
Xuecheng Liu, Daman Arora, Gokul Swamy +1
While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is a computational bottleneck.…
cs.CR2026
Image Prompt Reconstruction Attacks on Distributed MLLM Inference Frameworks
Xinjian Luo, Hongyan Chang, Jianxin Wei +5
Distributed large language model (LLM) inference frameworks connect isolated consumer-grade devices for large-scale model inference, substantially reducing hardware constraints. Ho…