collaborators

5 papers

cs.LG2026

Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution

Mingyu Yang, Keye Zheng, Congchao Cheng +4

In batch-style trace distillation, the same memory operation may receive contradictory feedback across different batches. Existing methods lack a cross-batch, operation-level evide…

cs.CL2026

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling

Parsa Ashrafi Fashi, Utkarsh Saxena, Mehdi Rezagholizadeh +7

Hybrid sequence models that combine efficient Transformer components with linear sequence modeling blocks are a promising alternative to pure Transformers, but most are still pretr…

cs.LG2026

UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience

Zichuan Lin, Feiyu Liu, Yijun Yang +9

Autonomous mobile GUI agents have attracted increasing attention along with the advancement of Multimodal Large Language Models (MLLMs). However, existing methods still suffer from…

cs.DC2025

DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving

Mingyu Yang, Jae-Young Choi, Kihyo Moon +2

Speculative decoding accelerates large language model inference, but its reliance on a fixed speculation length is suboptimal in large-batch serving environments with diverse reque…

cs.DC2025

ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor

Seungbeom Choi, Jeonghoe Goo, Eunjoo Jeon +2

We propose ELIS, a serving system for Large Language Models (LLMs) featuring an Iterative Shortest Remaining Time First (ISRTF) scheduler designed to efficiently manage inference t…