2 papers
cs.CL2025
A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
Fali Wang, Jihai Chen, Shuhua Yang +4
Large language models (LLMs) have achieved remarkable progress across domains and applications but face challenges such as high fine-tuning costs, inference latency, limited edge d…
cs.LG2025
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
Fali Wang, Jihai Chen, Shuhua Yang +7
Test-Time Scaling (TTS) improves large language models (LLMs) by allocating additional computation during inference, typically through parallel, sequential, or hybrid scaling. Howe…