3 papers
cs.CL2026
Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming
Pride Kavumba, Koki Wataoka, Huy H. Nguyen +2
In many practical LLM deployments, a single guardrail is used for both prompt and response moderation. Prompt moderation operates on fully observed text, whereas streaming response…
cs.CL2025
Self-Preference Bias in LLM-as-a-Judge
Koki Wataoka, Tsubasa Takahashi, Ryokan Ri
Automated evaluation leveraging large language models (LLMs), commonly referred to as LLM evaluators or LLM-as-a-judge, has been widely used in measuring the performance of dialogu…
cs.CR2025
MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language Models
Shojiro Yamabe, Futa Waseda, Tsubasa Takahashi +1
Protecting the intellectual property of Large Language Models (LLMs) has become increasingly critical due to the high cost of training. Model merging, which integrates multiple exp…