24 papers
When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL
Jiaqian Li
Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned transformations and attention r…
HACO: Hedged Agent Computing for Reliable LLM Systems
Enhan Li, Hongyang Du
As large language model (LLM) agents move from isolated prompting to longhorizon workflows, failures increasingly arise at the role-to-instance binding boundary, where task-specifi…
MORES: Mobile Reasoning-as-a-Service via Distributed LLM Inference-Time Scaling
Guanchen Liu, Hongyang Du, Kaibin Huang
Inference-time scaling has emerged as an effective approach for enhancing the capabilities of Large Language Models (LLMs), addressing the growing demand for stronger reasoning wit…
Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge
Haotian Zheng, Zhanwei Wang, Mingyao Cui +3
Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment t…
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
Hongyang Du, Junjie Ye, Xiaoyan Cong +7
While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deforma…
ReaCritic: Reasoning Transformer-based DRL Critic-model Scaling For Wireless Networks
Feiran You, Hongyang Du
Heterogeneous Networks (HetNets) pose critical challenges for intelligent management due to the diverse user requirements and time-varying wireless conditions. These factors introd…