2 citations · 3 across the 25 of their papers we have counts for
9 papers · 1 filter
NebulaSD: Many-for-Many Speculative Decoding
Junhao He, Hongyang Du
Speculative decoding accelerates Large Language Model (LLM) inference by using a lightweight draft model to propose candidate tokens for parallel verification by a target model. Dr…
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Hongyang Du, Lan Yan, Christian Flores +1
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programma…
CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
Enhan Li, Junhao He, Hongyang Du
On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervisi…
HACO: Hedged Agent Computing for Reliable LLM Systems
Enhan Li, Hongyang Du
As large language model (LLM) agents move from isolated prompting to longhorizon workflows, failures increasingly arise at the role-to-instance binding boundary, where task-specifi…
MORES: Mobile Reasoning-as-a-Service via Distributed LLM Inference-Time Scaling
Guanchen Liu, Hongyang Du, Kaibin Huang
Inference-time scaling has emerged as an effective approach for enhancing the capabilities of Large Language Models (LLMs), addressing the growing demand for stronger reasoning wit…
Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge
Haotian Zheng, Zhanwei Wang, Mingyao Cui +3
Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment t…