Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Data Contamination Can Cross Language Barriers
Feng Yao, Yufan Zhuang, Zihao Sun +3
The opacity in developing large language models (LLMs) is raising growing concerns about the potential contamination of public benchmarks in the pre-training data. Existing contami…
cs.CL2024
Beyond Scaling: Predicting Patent Approval with Domain-specific Fine-grained Claim Dependency Graph
Xiaochen Kev Gao, Feng Yao, Kewen Zhao +4
Model scaling is becoming the default choice for many language tasks due to the success of large language models (LLMs). However, it can fall short in specific scenarios where simp…