Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring
Yang Gao
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety…
cs.CL2026
Aligning Large Language Models with Searcher Preferences
Wei Wu, Peilun Zhou, Liyi Chen +6
The paradigm shift from item-centric ranking to answer-centric synthesis is redefining the role of search engines. While recent industrial progress has applied generative technique…
cs.CL2025
FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models
Yan Gao, Massimo Roberto Scamarcia, Javier Fernandez-Marques +18
Large Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raisin…