3 papers
cs.LG2026
Online Scheduling for LLM Inference with KV Cache Constraints
Patrick Jaillet, Jiashuo Jiang, Konstantina Mellou +3
Large Language Model (LLM) inference, where a trained model generates text one word at a time in response to user prompts, is a computationally intensive process requiring efficien…
cs.AI2025
MPrune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation
Weizi Shao, Taolin Zhang, Zijie Zhou +3
Recent advancements in multi-modal retrieval-augmented generation (mRAG), which enhance multi-modal large language models (MLLMs) with external knowledge, have demonstrated that th…
cs.GT2025
Grace Period is All You Need: Individual Fairness without Revenue Loss in Revenue Management
Patrick Jaillet, Chara Podimata, Zijie Zhou
Imagine you and a friend purchase identical items at a store, yet only your friend received a discount. Would your friend's discount make you feel unfairly treated by the store? An…