2 citations · 2 across the 5 of their papers we have counts for
6 papers
TPO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning
Haixin Wang, Hejie Cui, Chenwei Zhang +7
Recent progress in multi-turn reinforcement learning (RL) has significantly improved reasoning LLMs' performances on complex interactive tasks. Despite advances in stabilization te…
MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search
Minkyoung Cho, Insu Jang, Shuowei Jin +5
Fine-tuning Multimodal Large Language Models (MLLMs) with parameter-efficient methods like Low-Rank Adaptation (LoRA) is crucial for task adaptation. However, imbalanced training d…
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
Yongji Wu, Xueshen Liu, Shuowei Jin +6
The Mixture-of-Experts (MoE) architecture has become increasingly popular as a method to scale up large language models (LLMs). To save costs, heterogeneity-aware training solution…
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
The AIBrix Team, Jiaxin Shan, Varun Gupta +24
We introduce AIBrix, a cloud-native, open-source framework designed to optimize and simplify large-scale LLM deployment in cloud environments. Unlike traditional cloud-native stack…
Eagle: Efficient Training-Free Router for Multi-LLM Inference
Zesen Zhao, Shuowei Jin, Z. Morley Mao
The proliferation of Large Language Models (LLMs) with varying capabilities and costs has created a need for efficient model selection in AI systems. LLM routers address this need…
Compute Or Load KV Cache? Why Not Both?
Shuowei Jin, Xueshen Liu, Qingzhao Zhang +1
Large Language Models (LLMs) are increasingly deployed in large-scale online services, enabling sophisticated applications. However, the computational overhead of generating key-va…