activity
20212026
most citedLongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

13 citations · 49 across the 39 of their papers we have counts for

collaborators
Showing 2024 · cs.CLShow all

6 papers · 2 filters

cs.CL2024★ 1 cited

SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Yucheng Li, Huiqiang Jiang, Qianhui Wu +8

Long-context LLMs have enabled numerous downstream applications but also introduced significant challenges related to computational and memory efficiency. To address these challeng…

cs.CL2024★ 3 cited

TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning

Shivam Shandilya, Menglin Xia, Supriyo Ghosh +4

The increasing prevalence of large language models (LLMs) such as GPT-4 in various applications has led to a surge in the size of prompts required for optimal performance, leading…

cs.CL2024★ 4 cited

MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Huiqiang Jiang, Yucheng Li, Chengruidong Zhang +9

The computational challenges of Large Language Model (LLM) inference remain a significant barrier to their widespread deployment, especially as prompt lengths continue to increase.…

cs.CL2024★ 2 cited

Mitigate Position Bias in Large Language Models via Scaling a Single Dimension

Yijiong Yu, Huiqiang Jiang, Xufang Luo +6

Large Language Models (LLMs) are increasingly applied in various real-world scenarios due to their excellent generalization capabilities and robust generative abilities. However, t…

cs.CL2024

Position Engineering: Boosting Large Language Models through Positional Information Manipulation

Zhiyuan He, Huiqiang Jiang, Zilong Wang +3

The performance of large language models (LLMs) is significantly influenced by the quality of the prompts provided. In response, researchers have developed enormous prompt engineer…

cs.CL2024★ 2 cited

LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression

Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang +10

This paper focuses on task-agnostic prompt compression for better generalizability and efficiency. Considering the redundancy in natural language, existing approaches compress prom…