4 papers
QuoteBench: How Matched Scores Can Hide Command-Path Failures
Shangao Li, Yao Zhang, Volker Tresp +1
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation er…
Exact Rank and Convex Calibration Dimension Lower Bounds for the Multi-Label F1 Loss
Mingyuan Zhang
The instance-wise measure is a central performance measure for multi-label classification. For a problem with labels, it defines a loss matrix. Previous w…
Multiclass Learning from Noisy Labels for Non-decomposable Performance Measures
Mingyuan Zhang, Shivani Agarwal
There has been much interest in recent years in learning good classifiers from data with noisy labels. Most work on learning from noisy labels has focused on standard loss-based pe…
On the Minimax Regret in Online Ranking with Top-k Feedback
Mingyuan Zhang, Ambuj Tewari
In online ranking, a learning algorithm sequentially ranks a set of items and receives feedback on its ranking in the form of relevance scores. Since obtaining relevance scores typ…