Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023
Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference
Junyan Li, Li Lyna Zhang, Jiahang Xu +9
Deploying pre-trained transformer models like BERT on downstream tasks in resource-constrained scenarios is challenging due to their high inference cost, which grows rapidly with i…
cs.CL2023
UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation
Daixuan Cheng, Shaohan Huang, Junyu Bi +7
Large Language Models (LLMs) are popular for their impressive abilities, but the need for model-specific fine-tuning or task-specific prompt engineering can hinder their generaliza…