1 paper
Ruiyan Qi, Congding Wen, Weibo Zhou +3
Evaluating large language models (LLMs) in specific domain like tourism remains challenging due to the prohibitive cost of annotated benchmarks and persistent issues like hallucina…