3 papers
cs.AI2026
Building Interpretable Models for Moral Decision-Making
Mayank Goel, Aritra Das, Paras Chopra
We build a custom transformer model to study how neural networks make moral decisions on trolley-style dilemmas. The model processes structured scenarios using embeddings that enco…
cs.CL2025
IPO: Your Language Model is Secretly a Preference Classifier
Shivank Garg, Ayush Singh, Shweta Singh +1
Reinforcement learning from human feedback (RLHF) has emerged as the primary method for aligning large language models (LLMs) with human preferences. While it enables LLMs to achie…
cs.AI2025
Do GFlowNets Transfer? Case Study on the Game of 24/42
Adesh Gupta, Abhinav Kumar, Mansi Gupta +1
Generating diverse solutions is key to human-like reasoning, yet autoregressive language models focus on single accurate responses, limiting creativity. GFlowNets optimize solution…