5 papers
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
Yuchen Yang, Yaru Zhao, Pu Yang +2
While Mixture-of-Experts (MoE) architectures substantially bolster the expressive power of large-language models, their prohibitive memory footprint severely impedes the practical…
SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
Xiaomeng Yang, Mengping Yang, Junyan Wang +3
Preference learning has garnered extensive attention as an effective technique for aligning diffusion models with human preferences in visual generation. However, existing alignmen…
Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization
Yuchen Shi, Yuzheng Cai, Siqi Cai +15
Existing Large Language Model (LLM) agent frameworks face two significant challenges: high configuration costs and static capabilities. Building a high-quality agent often requires…
From Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMs
Jiaxiang Chen, Zhuo Wang, Mingxi Zou +4
Large language models (LLMs) have advanced general-purpose reasoning, showing strong performance across diverse tasks. However, existing methods often rely on implicit exploration,…
Get Experience from Practice: LLM Agents with Record & Replay
Erhu Feng, Wenbo Zhou, Zibin Liu +8
AI agents, empowered by Large Language Models (LLMs) and communication protocols such as MCP and A2A, have rapidly evolved from simple chatbots to autonomous entities capable of ex…