3 papers
cs.CV2026
AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty
Yan Ma, Lizhuo Zhang
Multimodal large language models (MLLMs) are widely used for automated annotation, yet their per-class accuracy varies widely (e.g., 12%-98% across the 13 classes of three classroo…
cs.CL2026
Decomposing Wrong-Consensus Agreement in LLM Self-Consistency
Lizhuo Zhang, Mengmeng Tang, Chenfeng Long +2
Agreement among repeated samples of a language model is routinely read as evidence about answer reliability, yet wrong answers can agree just as strongly as right ones. This paper…
cs.LG2026
Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets
Yan Ma, Lizhuo Zhang
Across seven public educational prediction datasets, three passed all four pre-modeling reliability checks; the remaining four either failed group-aware generalization tests or lac…