3 papers
cs.CR2026
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
Doguhuan Yeke, Yanming Zhou, Leo Y. Lin +3
Recent advances in Vision-Language Models (VLMs) facilitate a new class of embodied AI systems, where these models are integrated into physical platforms, e.g. robots and autonomou…
cs.CL2026
JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
Leonard Lin, Adam Lensenmayer
We introduce JP-TL-Bench, a lightweight, open benchmark designed to guide the iterative development of Japanese-English translation systems. In this context, the challenge is often…
cs.CL2025
WebNovelBench: Placing LLM Novelists on the Web Novel Distribution
Leon Lin, Jun Zheng, Haidong Wang
Robustly evaluating the long-form storytelling capabilities of Large Language Models (LLMs) remains a significant challenge, as existing benchmarks often lack the necessary scale,…