3 papers
cs.LG2025
Scaling Unverifiable Rewards: A Case Study on Visual Insights
Shuyu Gan, James Mooney, Pan Hao +4
Large Language Model (LLM) agents can increasingly automate complex reasoning through Test-Time Scaling (TTS), iterative refinement guided by reward signals. However, many real-wor…
cs.LG2025
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
Ioannis Tsaknakis, Bingqing Song, Shuyu Gan +5
Large Language Models (LLMs) excel at producing broadly relevant text, but this generality becomes a limitation when user-specific preferences are required, such as recommending re…
cs.CL2025
Effectively Steer LLM To Follow Preference via Building Confident Directions
Bingqing Song, Boran Han, Shuai Zhang +5
Having an LLM that aligns with human preferences is essential for accommodating individual needs, such as maintaining writing style or generating specific topics of interest. The m…