2 papers
cs.CL2024
PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models
Huixuan Zhang, Yun Lin, Xiaojun Wan
Large language models (LLMs) are known to be trained on vast amounts of data, which may unintentionally or intentionally include data from commonly used benchmarks. This inclusion…
cs.IR2024
Automated Similarity Metric Generation for Recommendation
Liang Qu, Yun Lin, Wei Yuan +3
The embedding-based architecture has become the dominant approach in modern recommender systems, mapping users and items into a compact vector space. It then employs predefined sim…