3 papers
cs.CL2026
Why Expert Alignment Is Hard: Evidence from Subjective Evaluation
Tzu-Mi Lin, Wataru Hirota, Tatsuya Ishigaki +2
Aligning large language models with expert judgment is especially difficult in subjective evaluation tasks, where experts may disagree, rely on tacit criteria, and change their jud…
cs.CL2026
Aggregate vs. Personalized Judges in Business Idea Evaluation: Evidence from Expert Disagreement
Wataru Hirota, Tomoki Taniguchi, Tomoko Ohkuma +6
Evaluating LLM-generated business ideas is often harder to scale than generating them. Unlike standard NLP benchmarks, business idea evaluation relies on multi-dimensional criteria…
cs.CL2025
Exploring Design of Multi-Agent LLM Dialogues for Research Ideation
Keisuke Ueda, Wataru Hirota, Takuto Asakura +4
Large language models (LLMs) are increasingly used to support creative tasks such as research idea generation. While recent work has shown that structured dialogues between LLMs ca…