1 paper
Jessica M. Lundin, Usman Nasir Nakakana, Guillaume Chabot-Couture
Rigorous evaluation of domain-specific language models requires benchmarks that are comprehensive, contamination-resistant, and maintainable. Static, manually curated datasets do n…