2 papers
cs.CL2024
Scaling LLM Inference with Optimized Sample Compute Allocation
Kexun Zhang, Shang Zhou, Danqing Wang +2
Sampling is a basic operation in many inference-time algorithms of large language models (LLMs). To scale up inference efficiently with a limited compute, it is crucial to find an…
cs.CL2024
Evaluating the Smooth Control of Attribute Intensity in Text Generation with LLMs
Shang Zhou, Feng Yao, Chengyu Dong +2
Controlling the attribute intensity of text generation is crucial across scenarios (e.g., writing conciseness, chatting emotion, and explanation clarity). The remarkable capabiliti…