activity
20242026
collaborators

5 papers

cs.LG2026

Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in

Feiyu Wang, Xinyu Tan, Bokai Huang +4

Large language models (LLMs) have revolutionized artificial intelligence, yet their massive memory and computational demands necessitate aggressive quantization, increasingly pushi…

cs.AI2025

LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation

Liya Zhu, Peizhuang Cong, Jingzhe Ding +17

Large Language Models (LLMs) perform well on standard reasoning and question-answering benchmarks, yet such evaluations often fail to capture their ability to handle long-tail, exp…

cs.LG2025

Rank Also Matters: Hierarchical Configuration for Mixture of Adapter Experts in LLM Fine-Tuning

Peizhuang Cong, Wenpu Liu, Wenhan Yu +2

Large language models (LLMs) have demonstrated remarkable success across various tasks, accompanied by a continuous increase in their parameter size. Parameter-efficient fine-tunin…

cs.LG2024

BATON: Enhancing Batch-wise Inference Efficiency for Large Language Models via Dynamic Re-batching

Peizhuang Cong, Qizhi Chen, Haochen Zhao +1

The advanced capabilities of Large Language Models (LLMs) have inspired the development of various interactive web services or applications, such as ChatGPT, which offer query infe…

cs.LG2024

INT-FlashAttention: Enabling Flash Attention for INT8 Quantization

Shimao Chen, Zirui Liu, Zhiying Wu +6

As the foundation of large language models (LLMs), self-attention module faces the challenge of quadratic time and memory complexity with respect to sequence length. FlashAttention…