4 papers
Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal Routing
Srinivasan Manoharan, Junhua Zhao, Fangbo Tu +6
Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries, escalations, and developer…
Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation
Fangbo Tu, Junhua Zhao, Chi Liu +4
Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost constrained production enviro…
Multi-Objective Large Language Model Unlearning
Zibin Pan, Shuwen Zhang, Yuesheng Zheng +3
Machine unlearning in the domain of large language models (LLMs) has attracted great attention recently, which aims to effectively eliminate undesirable behaviors from LLMs without…
Federated Unlearning with Gradient Descent and Conflict Mitigation
Zibin Pan, Zhichao Wang, Chi Li +4
Federated Learning (FL) has received much attention in recent years. However, although clients are not required to share their data in FL, the global model itself can implicitly re…