4 papers
OdysSim: Building Foundation Models for Human Behavior Simulation
Xuhui Zhou, Weiwei Sun, Weihua Du +6
Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homog…
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
Xiaoyan Bai, Alexander Baumgartner, Haojia Sun +2
Reproducibility crises across sciences highlight the limitations of the paper-centric review system in assessing the rigor and reproducibility of research. AI agents that autonomou…
LLM as Explainable Re-Ranker for Recommendation System
Yaqi Wang, Haojia Sun, Shuting Zhang
The application of large language models (LLMs) in recommendation systems has recently gained traction. Traditional recommendation systems often lack explainability and suffer from…
Retrieval-Augmented Generation for Domain-Specific Question Answering: A Case Study on Pittsburgh and CMU
Haojia Sun, Yaqi Wang, Shuting Zhang
We designed a Retrieval-Augmented Generation (RAG) system to provide large language models with relevant documents for answering domain-specific questions about Pittsburgh and Carn…