30 citations · 32 across the 7 of their papers we have counts for
7 papers
OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning
Jikai Long, Zijian Hu, Xiaodong Yu +2
Fine-tuning large language models (LLMs) using zeroth-order optimization (ZO) offers a memory-efficient alternative to gradient-based methods but suffers from slower convergence an…
Alopex: A Computational Framework for Enabling On-Device Function Calls with LLMs
Yide Ran, Zhaozhuo Xu, Yuhang Yao +9
The rapid advancement of Large Language Models (LLMs) has led to their increased integration into mobile devices for personalized assistance, which enables LLMs to call external AP…
Fox-1: Open Small Language Model for Cloud and Edge
Zijian Hu, Jipeng Zhang, Rui Pan +9
We present Fox-1, a series of small language models (SLMs) consisting of Fox-1-1.6B and Fox-1-1.6B-Instruct-v0.1. These models are pre-trained on 3 trillion tokens of web-scraped d…
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
Yuhang Yao, Han Jin, Alay Dilipbhai Shah +7
Large language models (LLMs) have surged in popularity and are extensively used in commercial applications, where the efficiency of model serving is crucial for the user experience…
TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
Dimitris Stripelis, Zijian Hu, Jipeng Zhang +6
With the rapid growth of Large Language Models (LLMs) across various domains, numerous new LLMs have emerged, each possessing domain-specific expertise. This proliferation has high…
TorchOpera: A Compound AI System for LLM Safety
Shanshan Han, Zijian Hu, Alay Dilipbhai Shah +5
We introduce TorchOpera, a compound AI system for enhancing the safety and quality of prompts and responses for Large Language Models. TorchOpera ensures that all user prompts are…