2 papers
cs.CL2025
A Systematic Analysis of Base Model Choice for Reward Modeling
Kian Ahrabian, Pegah Jandaghi, Negar Mokhberian +2
Reinforcement learning from human feedback (RLHF) and, at its core, reward modeling have become a crucial part of training powerful large language models (LLMs). One commonly overl…
cs.CR2025
Trim My View: An LLM-Based Code Query System for Module Retrieval in Robotic Firmware
Sima Arasteh, Pegah Jandaghi, Nicolaas Weideman +4
The software compilation process has a tendency to obscure the original design of the system and makes it difficult both to identify individual components and discern their purpose…