5 papers
RLHF May Not Reflect Genuine Preferences
Bijean Ghafouri, Eun Cheol Choi, Priyanka Dey +1
Reinforcement Learning from Human Feedback (RLHF) assumes that annotation responses reflect genuine human preferences. They often do not. Behavioral scientists have documented for…
What do people want to fact-check?
Bijean Ghafouri, Dorsaf Sallami, Luca Luceri +4
Research on misinformation has focused almost exclusively on supply, asking what falsehoods circulate, who produces them, and whether corrections work. A basic demand-side question…
The Variance Paradox: How AI Reduces Diversity but Increases Novelty
Bijean Ghafouri
The diversity of human expression is the raw material of discovery. Generative artificial intelligence threatens this resource even as it promises to accelerate innovation, a parad…
Lost Before Translation: Social Information Transmission and Survival in AI-AI Communication
Bijean Ghafouri, Emilio Ferrara
When AI systems summarize and relay information, they inevitably transform it. But how? We introduce an experimental paradigm based on the telephone game to study what happens when…
Epistemic Integrity in Large Language Models
Bijean Ghafouri, Shahrad Mohammadzadeh, James Zhou +8
Large language models are increasingly relied upon as sources of information, but their propensity for generating false or misleading statements with high confidence poses risks fo…