3 citations · 3 across the 12 of their papers we have counts for
13 papers
Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation
Seth Grief-Albert, Jessica Bo, Difan Jiao +1
Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, wh…
Synthetic Persona Pretraining: Alignment from Token Zero
Julian Minder, Viktor Moskvoretskii, Raghav Singhal +12
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant…
Tandem Reinforcement Learning with Verifiable Rewards
Difan Jiao, Raghav Singhal, Robert West +1
Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reaching expert or even superhuman performance i…
SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents
Qianfeng Wen, Yifan Simon Liu, Xin Liu +4
Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems. In recommendation agents, this creates a risk that…
MINER: Mining Multimodal Internal Representation for Efficient Retrieval
Weien Li, Rui Song, Zeyu Li +8
Visual document retrieval has become essential for accessing information in visually rich documents. Existing approaches fall into two camps. Late-interaction retrievers achieve st…
LLM Safety From Within: Detecting Harmful Content with Internal Representations
Difan Jiao, Yilun Liu, Ye Yuan +4
Guard models are widely used to detect harmful content in user prompts and LLM responses. However, state-of-the-art guard models rely solely on terminal-layer representations and o…