activity
20222024
most citedLongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

15 citations · 19 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL202415 cited

LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Yiran Ding, Li Lyna Zhang, Chengruidong Zhang +5

Large context window is a desirable feature in large language models (LLMs). However, due to high fine-tuning costs, scarcity of long texts, and catastrophic values introduced by n…

cs.AI20231 cited

Compresso: Structured Pruning with Collaborative Prompting Learns Compact Large Language Models

Song Guo, Jiahang Xu, Li Lyna Zhang +1

Despite the remarkable success of Large Language Models (LLMs), the massive size poses significant deployment challenges, particularly on resource-constrained hardware. While exist…

cs.CL2023

Accurate and Structured Pruning for Efficient Automatic Speech Recognition

Huiqiang Jiang, Li Lyna Zhang, Yuang Li +7

Automatic Speech Recognition (ASR) has seen remarkable advancements with deep neural networks, such as Transformer and Conformer. However, these models typically have large model s…

cs.CV2023

ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile Devices

Chen Tang, Li Lyna Zhang, Huiqiang Jiang +6

Neural Architecture Search (NAS) has shown promising performance in the automatic design of vision transformers (ViT) exceeding 1G FLOPs. However, designing lightweight and low-lat…

cs.CV20232 cited

SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 Inference

Li Lyna Zhang, Xudong Wang, Jiahang Xu +6

The combination of Neural Architecture Search (NAS) and quantization has proven successful in automatically designing low-FLOPs INT8 quantized neural networks (QNN). However, direc…

cs.IR20221 cited

SwiftPruner: Reinforced Evolutionary Pruning for Efficient Ad Relevance

Li Lyna Zhang, Youkow Homma, Yujing Wang +5

Ad relevance modeling plays a critical role in online advertising systems including Microsoft Bing. To leverage powerful transformers like BERT in this low-latency setting, many ex…