activity
20232026
most citedRenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

10 citations · 15 across the 16 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.CL2024

CORD: Balancing COnsistency and Rank Distillation for Robust Retrieval-Augmented Generation

Youngwon Lee, Seung-won Hwang, Daniel Campos +3

With the adoption of retrieval-augmented generation (RAG), large language models (LLMs) are expected to ground their generation to the retrieved contexts. Yet, this is hindered by…

cs.CL2024

Inference Scaling for Bridging Retrieval and Augmented Generation

Youngwon Lee, Seung-won Hwang, Daniel Campos +3

Retrieval-augmented generation (RAG) has emerged as a popular approach to steering the output of a large language model (LLM) by incorporating retrieved contexts as inputs. However…

cs.LG2024

SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation

Aurick Qiao, Zhewei Yao, Samyam Rajbhandari +1

LLM inference for enterprise applications, such as summarization, RAG, and code-generation, typically observe much longer prompt than generations, leading to high prefill cost and…

cs.LG2024

STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning

Jaeseong Lee, seung-won hwang, Aurick Qiao +3

Mixture-of-experts (MoEs) have been adopted for reducing inference costs by sparsely activating experts in Large language models (LLMs). Despite this reduction, the massive number…

cs.CL2024

Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding

Zhenyu Zhang, Runjin Chen, Shiwei Liu +5

This paper aims to overcome the "lost-in-the-middle" challenge of large language models (LLMs). While recent advancements have successfully enabled LLMs to perform stable language…

cs.LG2024★ 3 cited

FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design

Haojun Xia, Zhen Zheng, Xiaoxia Wu +10

Six-bit quantization (FP6) can effectively reduce the size of large language models (LLMs) and preserve the model quality consistently across varied applications. However, existing…