7 papers
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
Langming Liu, Kangtao Lv, Haibin Chen +8
Large language models (LLMs), despite their powerful capabilities, suffer from factual hallucinations where they generate verifiable falsehoods. We identify a root of this issue: t…
Data Distribution Matters: A Data-Centric Perspective on Context Compression for Large Language Model
Kangtao Lv, Jiwei Tang, Langming Liu +7
The deployment of Large Language Models (LLMs) in long-context scenarios is hindered by computational inefficiency and significant information redundancy. Although recent advanceme…
How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
Kangtao Lv, Haibin Chen, Yujin Yuan +5
Large language models (LLMs) have attracted significant attention due to their impressive general capabilities across diverse downstream tasks. However, without domain-specific opt…
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
Chongjie Si, Kangtao Lv, Jingjing Jiang +6
Model merging offers a training-free alternative to multi-task learning by combining independently fine-tuned models into a unified one without access to raw data. However, existin…
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
Haibin Chen, Kangtao Lv, Chengwei Hu +8
With the increasing use of Large Language Models (LLMs) in fields such as e-commerce, domain-specific concept evaluation benchmarks are crucial for assessing their domain capabilit…
HyperDet: Generalizable Detection of Synthesized Images by Generating and Merging A Mixture of Hyper LoRAs
Huangsen Cao, Yongwei Wang, Yinfeng Liu +6
The emergence of diverse generative vision models has recently enabled the synthesis of visually realistic images, underscoring the critical need for effectively detecting these ge…