Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
Leonard Lin, Adam Lensenmayer
We introduce JP-TL-Bench, a lightweight, open benchmark designed to guide the iterative development of Japanese-English translation systems. In this context, the challenge is often…
cs.CL2025
WebNovelBench: Placing LLM Novelists on the Web Novel Distribution
Leon Lin, Jun Zheng, Haidong Wang
Robustly evaluating the long-form storytelling capabilities of Large Language Models (LLMs) remains a significant challenge, as existing benchmarks often lack the necessary scale,…
cs.CL2024
Rethinking How to Evaluate Language Model Jailbreak
Hongyu Cai, Arjun Arunasalam, Leo Y. Lin +2
Large language models (LLMs) have become increasingly integrated with various applications. To ensure that LLMs do not generate unsafe responses, they are aligned with safeguards t…