1 paper
Artem Karpov, Seong Hah Cho, Austin Meek +3
In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same…