5 papers
Why Expert Alignment Is Hard: Evidence from Subjective Evaluation
Tzu-Mi Lin, Wataru Hirota, Tatsuya Ishigaki +2
Aligning large language models with expert judgment is especially difficult in subjective evaluation tasks, where experts may disagree, rely on tacit criteria, and change their jud…
Aggregate vs. Personalized Judges in Business Idea Evaluation: Evidence from Expert Disagreement
Wataru Hirota, Tomoki Taniguchi, Tomoko Ohkuma +6
Evaluating LLM-generated business ideas is often harder to scale than generating them. Unlike standard NLP benchmarks, business idea evaluation relies on multi-dimensional criteria…
Exploring Design of Multi-Agent LLM Dialogues for Research Ideation
Keisuke Ueda, Wataru Hirota, Takuto Asakura +4
Large language models (LLMs) are increasingly used to support creative tasks such as research idea generation. While recent work has shown that structured dialogues between LLMs ca…
Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
Kosuke Takahashi, Takahiro Omi, Kosuke Arima +1
The development of Large Language Models (LLMs) in various languages has been advancing, but the combination of non-English languages with domain-specific contexts remains underexp…
Training Generative Question-Answering on Synthetic Data Obtained from an Instruct-tuned Model
Kosuke Takahashi, Takahiro Omi, Kosuke Arima +1
This paper presents a simple and cost-effective method for synthesizing data to train question-answering systems. For training, fine-tuning GPT models is a common practice in resou…