4 papers
Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence
Zhen Yang, Hongyi Lin, Yifan He +7
In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread adoption of Pretrained Language Models…
In-depth Analysis on Caching and Pre-fetching in Mixture of Experts Offloading
Shuning Lin, Yifan He, Yitong Chen
In today's landscape, Mixture of Experts (MoE) is a crucial architecture that has been used by many of the most advanced models. One of the major challenges of MoE models is that t…
Cash or Comfort? How LLMs Value Your Inconvenience
Mateusz Cedro, Timour Ichmoukhamedov, Sofie Goethals +3
Large Language Models (LLMs) are increasingly proposed as near-autonomous artificial intelligence (AI) agents capable of making everyday decisions on behalf of humans. Although LLM…
Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations Categories
Tianlong Wang, Xianfeng Jiao, Yinghao Zhu +6
Recent studies have indicated that Large Language Models (LLMs) harbor an inherent understanding of truthfulness, yet often fail to consistently express it and generate false state…