3 papers
cs.CL2026
Dimensionality in Satisfaction Ratings
Andrew Hong, Jason Potteiger
We used a large language model (GPT-4.1) to annotate the text of about 9,000 support conversations at a global consumer-goods firm, decomposing customer-care satisfaction into comp…
cs.CL2026
The signal is the ceiling: Measurement limits of LLM-predicted experience ratings from open-ended survey text
Andrew Hong, Jason Potteiger, Luis E. Zapata
An earlier paper (Hong, Potteiger, and Zapata 2026) established that an unoptimized GPT 4.1 prompt predicts fan-reported experience ratings within one point 67% of the time from op…
cs.CL2026
LLM Predictive Scoring and Validation: Inferring Experience Ratings from Unstructured Text
Jason Potteiger, Andrew Hong, Ito Zapata
We tasked GPT-4.1 to read what baseball fans wrote about their game-day experience and predict the overall experience rating each fan gave on a 0-10 survey scale. The model receive…