2 papers
cs.CL2026
From Chains to Trees: Parent-Conditioned Drafting for Semi-Autoregressive Speculative Decoding
Zixian Li, Tong Li, Chi Xie +2
Speculative decoding accelerates LLM inference only when drafted continuations survive target-model verification. Semi-autoregressive drafters such as DSpark predict an entire toke…
cs.CV2025
AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model
Zhiwei Jin, Xiaohui Song, Nan Wang +36
In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hu…