activity
20242026
collaborators

5 papers

cs.CL2026

Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance

Robert Morabito, Tyler McDonald, Charitra Viswanath +4

Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away with markedly different rat…

cs.CL2025

Trace-of-Thought Prompting: Investigating Prompt-Based Knowledge Distillation Through Question Decomposition

Tyler McDonald, Ali Emami

Knowledge distillation allows smaller neural networks to emulate the performance of larger, teacher models with reduced computational demands. Traditional methods for Large Languag…

cs.CL2025

NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers

Angel Yahir Loredo Lopez, Tyler McDonald, Ali Emami

Large Language Models (LLMs) have shown impressive performance on various benchmarks, yet their ability to engage in deliberate reasoning remains questionable. We present NYT-Conne…

cs.CL2025

STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions

Robert Morabito, Sangmitra Madhusudan, Tyler McDonald +1

Mitigating explicit and implicit biases in Large Language Models (LLMs) has become a critical focus in the field of natural language processing. However, many current methodologies…

cs.CL2024

Can We Afford The Perfect Prompt? Balancing Cost and Accuracy with the Economical Prompting Index

Tyler McDonald, Anthony Colosimo, Yifeng Li +1

As prompt engineering research rapidly evolves, evaluations beyond accuracy are crucial for developing cost-effective techniques. We present the Economical Prompting Index (EPI), a…