3 papers
cs.AI2026
PROTEUS: SLA-Aware Routing via Lagrangian RL for Multi-LLM Serving Systems
Amit Singh Bhatti, Vishal Vaddina, Dagnachew Birru
Production LLM deployments serve diverse workloads where cost and quality requirements vary by customer tier, time of day, and query criticality. Model serving systems accept laten…
cs.AI2026
Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward
Gourab K Patro, Himanshi Agrawal, Himanshu Gharat +4
Modern general-purpose AI systems made using large language and vision models, are capable of performing a range of tasks like writing text articles, generating and debugging codes…
cs.LG2025
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
Aasheesh Singh, Vishal Vaddina, Dagnachew Birru
We introduce ORPO-Distill, a general-purpose method for cross-architecture LLM distillation that formulates the problem as a preference optimization task. Unlike standard CoT disti…