5 papers
LLM Performance Predictors: Learning When to Escalate in Hybrid Human-AI Moderation Systems
Or Bachar, Or Levi, Sardhendu Mishra +6
As LLMs are increasingly integrated into human-in-the-loop content moderation systems, a central challenge is deciding when their outputs can be trusted versus when escalation for…
Evaluating LLM Metrics Through Real-World Capabilities
Justin K Miller, Wenjia Tang
As generative AI becomes increasingly embedded in everyday workflows, it is important to evaluate its performance in ways that reflect real-world usage rather than abstract notions…
Balancing Complexity and Informativeness in LLM-Based Clustering: Finding the Goldilocks Zone
Justin Miller, Tristram Alexander
The challenge of clustering short text data lies in balancing informativeness with interpretability. Traditional evaluation metrics often overlook this trade-off. Inspired by lingu…
Moving Past Single Metrics: Exploring Short-Text Clustering Across Multiple Resolutions
Justin Miller, Tristram Alexander
Cluster number is typically a parameter selected at the outset in clustering problems, and while impactful, the choice can often be difficult to justify. Inspired by bioinformatics…
Human-interpretable clustering of short-text using large language models
Justin K. Miller, Tristram J. Alexander
Clustering short text is a difficult problem, due to the low word co-occurrence between short text documents. This work shows that large language models (LLMs) can overcome the lim…